S Assistant Box is an experimental lifelong-learning environment for small language models.
The project places a newly trained model, such as a 124M-parameter model, inside a simulated life. Instead of exposing the model to one enormous static dataset and expecting it to learn everything at once, the Box teaches the model progressively through structured interaction, feedback, practice, memory, and reflection.
The long-term goal is to study whether a small model can become substantially more useful through a cognitive architecture that provides:
- staged education;
- long-term episodic and semantic memory;
- skill and procedure storage;
- active practice and feedback;
- knowledge consolidation;
- motivation and goal selection;
- gradual specialization.
The project is a research experiment, not a claim that a 124M model will automatically become equivalent to a frontier model. The model will remain a 124M-parameter model. The complete system may become more capable because it combines the model with memory, tools, environment state, and learning mechanisms.
The central question is:
Can structured lifelong teaching, active experience, and memory consolidation give a small language model better sample efficiency, retention, transfer, and continual-learning behavior than conventional static training under a comparable compute budget?
This question is testable. The project must compare controlled baselines rather than rely on subjective impressions.
The Box represents a compressed, accelerated life cycle:
Infancy
-> language, symbols, basic concepts
Childhood
-> vocabulary, grammar, arithmetic, social rules
School
-> mathematics, science, reading, writing, problem solving
Adolescence
-> abstraction, debate, projects, independent learning
Adulthood
-> specialization, work, research, long-term goals
Reflection
-> review, consolidation, identity, and future planning
Logical time can be accelerated. A simulated year does not need to take one real year.
The initial learner should be a 124M-parameter causal language model that has completed general language pretraining but has not yet been instruction-tuned. This gives the Box a real language foundation while leaving instruction following, domain education, and lifelong adaptation for the Box itself.
The recommended first training sequence is:
FineWeb10B
-> causal language-model pretraining
-> checkpoint: 124M Base Model
-> HellaSwag evaluation
-> import checkpoint into S Assistant Box
-> Box-led instruction teaching and lifelong learning
FineWeb10B is appropriate as a large-scale pretraining corpus. HellaSwag
is primarily a multiple-choice commonsense-reasoning benchmark, so it should be
kept as a held-out evaluation set for the base model. Its training split may be
used later as an explicitly logged Box lesson or instruction-tuning experiment,
but it should not be mixed into the initial pretraining result if we want a
clean measurement of the Box's contribution.
Every training run should preserve:
- dataset versions and hashes;
- tokenizer files and vocabulary configuration;
- model configuration;
- random seeds;
- validation loss curves;
- HellaSwag accuracy;
- checkpoint and optimizer state;
- hardware and precision settings.
The parameter count remains 124M after pretraining, instruction tuning, or adapter updates. The complete Box system may have greater effective capacity because it also includes memory, tools, and environment state.
The Box may search the web when it needs high-quality teaching material or when it detects a knowledge gap. Search results must be treated as sources to verify, not as automatically trusted truth.
Search
-> Fetch source
-> Extract content
-> Rank and filter
-> Verify claims
-> Adapt to life stage
-> Teach lesson
-> Assign exercise
-> Assess and correct
-> Store source and learning evidence
The search subsystem should record URLs, publication dates, source text or snapshots where permitted, extracted claims, citations, and retrieval times. This makes dynamic web-based education auditable and reproducible. Search should be invoked selectively rather than on every interaction.
Web content must be converted into more than a retrieved paragraph. A complete lesson should contain an explanation, examples, exercises, feedback, and later review. Otherwise the system is only retrieval-augmented generation, not lifelong teaching.
The first implementation may use a deterministic search-and-extraction controller. A larger teacher model can be added later, but its outputs and compute cost must be logged so that the teacher's capabilities are not confused with the 124M learner's capabilities.
- Separate parameters from memory. The model weights, external memory, summaries, and learned adapters must be tracked separately.
- Teach progressively. The curriculum should follow prerequisites and increase in difficulty.
- Require active use. A fact is not considered learned merely because it was inserted into a database.
- Measure transfer. The model must solve new problems, not only repeat lessons.
- Record provenance. Every lesson, teacher response, feedback event, and memory update must be auditable.
- Use controlled experiments. Results must be compared with equivalent baselines and compute budgets.
- Avoid unsupported anthropomorphism. The system may simulate memory, motivation, and reflection without proving consciousness or human-like subjective experience.
S Assistant Box
├── Base Model
├── Life Environment
├── Role System
├── Curriculum Engine
├── Cognitive Loop
├── Memory System
├── Learning System
├── Motivation System
├── Evaluation System
└── User Interface
The initial implementation should support a small causal language model through standard Python model interfaces. The first target is a 124M-parameter model, but the system should not hard-code that size.
Responsibilities:
- language understanding and generation;
- basic reasoning;
- optional hidden-state extraction;
- optional adapter or LoRA updates.
The environment manages simulated time, events, tasks, resources, relationships, and consequences.
Example events include:
- lessons and conversations;
- reading and practice;
- examinations;
- mistakes and corrections;
- social interactions;
- career choices;
- projects and research tasks;
- periodic reviews.
Roles provide different forms of teaching and feedback:
Parent: language, safety, routines, and basic social behavior;Teacher: structured lessons, explanations, exercises, and grading;Friend: conversation and informal practice;Mentor: long-term guidance and specialization;Employer: practical tasks and performance feedback;Researcher: open-ended investigation and hypothesis testing.
The first version can use deterministic rules and templates. A larger teacher model may be used later, but its outputs and cost must be recorded so that the source of capability remains clear.
The curriculum controls what is taught and when. It should include:
- prerequisite graphs;
- difficulty progression;
- spaced review;
- adaptive exercise selection;
- error diagnosis;
- mastery tests;
- transfer tasks;
- stage transitions.
Each interaction follows this general loop:
Observe
-> Interpret
-> Retrieve relevant memory
-> Think and plan
-> Respond or act
-> Receive feedback
-> Store the episode
-> Consolidate knowledge when scheduled
The event log should capture the observation, retrieved memories, action, feedback, confidence, and later reuse of the knowledge.
The memory system is inspired by cognitive science, but it is not a literal simulation of a biological brain.
Short-lived context for the current task or conversation.
Specific experiences, including what happened, the model's response, the result, and the feedback.
Generalized facts and concepts extracted from many episodes.
Reusable skills and action sequences, such as solving an equation or writing an experiment report.
Recommended initial storage:
- SQLite for events and metadata;
- full-text search for exact terms;
- a vector index for semantic retrieval;
- a knowledge graph for concept relationships;
- summaries for stage-level consolidation.
The first implementation may use SQLite with FTS5 and a local vector index. A separate embedding model can be evaluated later; its contribution must be documented.
The current kernel separates lesson episodes from extracted semantic facts.
Simple declarative claims are extracted deterministically and stored with
source/provenance, confidence, and an explicit verified flag. Extraction is
not verification: a fact becomes verified only through an explicit evidence
event, so the Box does not silently treat every teacher sentence as truth.
Learning should operate at three levels:
- Memory write: store the raw experience.
- Memory consolidation: summarize, merge, link, and verify knowledge.
- Parameter adaptation: optionally update a small adapter or LoRA module.
This allows experiments with:
- external memory only;
- memory plus consolidation;
- memory plus parameter adaptation;
- all mechanisms combined.
The first parameter-adaptation implementation is an explicit online LoRA SFT path for nanochat d4 checkpoints. It freezes the base model, injects low-rank adapters into attention/MLP projections, trains only those adapters on prompt/answer pairs, and saves a small adapter checkpoint separately from the base weights. This makes it possible to measure whether a fact survives when the external memory is disabled. SFT labels use the causal next-token shift; plain Base checkpoints also receive a BOS terminator so answers do not run indefinitely or repeat after the target sentence.
Later versions may provide explicit drives and goals, such as:
- curiosity;
- uncertainty reduction;
- mastery;
- social reward;
- completion of long-term projects;
- preference formation.
These mechanisms create an agent with persistent objectives, but they do not by themselves demonstrate consciousness.
Python is the recommended language for the research prototype because it has the strongest ecosystem for model experimentation:
- PyTorch;
- Hugging Face Transformers and Datasets;
- NumPy;
- FastAPI;
- SQLite;
- FAISS or hnswlib;
- Pydantic.
Rust can optimize performance-critical components after the research design is validated:
- event scheduling;
- high-throughput retrieval;
- concurrent environment simulation;
- large event-log processing;
- long-running services.
Python can call Rust components through PyO3.
TypeScript is appropriate for a web interface containing:
- a conversation panel;
- a life timeline;
- a memory browser;
- a teacher console;
- evaluation dashboards;
- system and resource monitoring.
Recommended division:
Python AI and research core
Rust performance-critical components
TypeScript web interface
The first experiment should not attempt to simulate all of human civilization. It should use a small, controlled artificial world.
Suggested content:
- 100 words and symbols;
- simple grammar;
- 20 mathematical concepts;
- 10 basic science concepts;
- approximately 1,000 interactive tasks;
- repeated lessons, mistakes, corrections, and tests.
Compare at least three systems:
Model A: 124M model without external memory
Model B: 124M model with retrieval memory
Model C: 124M model with retrieval and consolidation
Later experiments can add adapter training and a larger curriculum.
The project should report measurable outcomes:
- factual recall;
- retention after delays;
- multi-turn consistency;
- error correction;
- transfer to unseen problems;
- reasoning and planning;
- forgetting and relearning;
- curriculum sample efficiency;
- CPU and memory cost;
- performance under a fixed compute budget.
The evaluation set must be separated from the teaching material. Repeating a lesson is not evidence of understanding.
Inference, retrieval, logging, and most memory operations can run on a CPU for a 124M model. Quantization and ONNX export may improve speed.
Adapter training and frequent parameter updates can also run on a CPU, but may be slow. The prototype should therefore begin with external memory and infrequent consolidation, then add parameter adaptation as an explicit experiment.
The runtime is hardware-agnostic:
device="auto"selects CUDA when available and otherwise uses CPU;- CUDA devices such as an RTX 3070 and an H100 use the same Box interfaces;
- the Box can run many life instances in parallel to improve GPU utilization;
- hardware details are exposed through the runtime status for experiment logs.
An H100 is excessive for one 124M model inference stream, but useful for parallel life simulations, teacher-model generation, adapter training, distillation, and large evaluation sweeps. An RTX 3070 is sufficient for the initial prototype and small-batch experiments.
s_assistant_box/
├── model/
├── environment/
├── curriculum/
├── roles/
├── cognition/
├── memory/
├── learning/
├── motivation/
├── evaluation/
├── api/
├── ui/
├── configs/
├── tests/
└── README.md
The first checkpoint produced by the training script can be loaded through the custom checkpoint adapter. From the repository root:
python scripts/test_checkpoint_in_box.py --checkpoint "./base model/log/model_000020.pt" --device cuda --dtype float16The test performs four observable actions:
- asks a baseline question before teaching;
- teaches three triangle lessons through the Box;
- asks the same question after teaching;
- prints retrieved memories, life status, and the event-log path.
This first test measures checkpoint integration and memory retrieval. It does not by itself prove parameter learning. Online SFT and LoRA updates are implemented separately and require memory-disabled, held-out, and reload evaluation. A meaningful Box learning result requires a stronger base checkpoint, a held-out question set, and controlled comparison with the retrieval and tool-assisted paths.
The d4 checkpoints produced by the local nanochat trainer are a different
architecture from the GPT-2-style 124M checkpoints. They use a 32K custom BPE
tokenizer, rotary attention, and value embeddings. The Box therefore loads the
original nanochat model implementation and tokenizer rather than attempting to
reinterpret the weights as GPT-2.
The adapter expects the nanochat source tree and tokenizer in the local
nanochat directory by default. Pass --nanochat-root path/to/nanochat when
the source tree is stored elsewhere. Run both the scientific Base Model and
the SFT usability baseline with:
python scripts/test_d4_checkpoints_in_box.py --device cuda --dtype float16 --max-new-tokens 64If CUDA is unavailable, use a CPU smoke test:
python scripts/test_d4_checkpoints_in_box.py --device cpu --dtype float32 --max-new-tokens 16The test uses an artificial fact (zorblite is violet) that is unlikely to be
present in pretraining data. It reports the baseline answer, three Box lessons,
the post-teaching answer, a memory-disabled control answer, the number of
retrieved memories, whether the retrieved memories contain the target fact, a
paraphrased delayed question, and a restarted-Box question.
Results are written to
box_runs/d4_box_report.json and one JSONL event log per checkpoint.
Interpret the result in separate layers:
- A successful load and generated response verifies checkpoint, architecture, and tokenizer compatibility.
- Retrieved memories verify that the Box recorded and recalled teaching.
- A correct answer on paraphrased and delayed held-out questions tests useful behavior, but can still be retrieval-augmented generation.
- To claim parameter learning, repeat the questions with memory disabled and after restarting the Box. The answer must remain correct without the stored lesson prompt. The included d4 online-SFT script can attach a LoRA adapter; its memory-disabled and held-out controls are required before claiming parameter adaptation.
The test also writes a portable semantic-memory snapshot for each checkpoint.
Loading that snapshot into a new MemoryStore verifies persistence separately
from model generation.
The default answer_mode="hybrid" keeps the raw model response in the event
log and uses a verified semantic fact as a clearly labeled fallback when the
model misses it. This improves user-facing answer quality without hiding the
raw-model metric. Use answer_mode="model" for a pure model measurement.
For durable long-running sessions, use the SQLite backend:
from s_assistant_box import SAssistantBox, SQLiteMemoryStore
memory = SQLiteMemoryStore("box_runs/lifelong_memory.sqlite3")
box = SAssistantBox(model, memory=memory)Call memory.close() when the process exits, or use it as a context manager.
To run online LoRA SFT on a d4 checkpoint:
python scripts/online_sft_d4.py --checkpoint "./d4/model_005140.pt" --device cuda --dtype float16 --steps 20 --rank 4The script prints the answer before training, trains only LoRA parameters,
saves box_runs/online_d4_lora.pt, and evaluates again with external memory
disabled. Start with 2-20 steps for a smoke test; longer runs can overfit the
small teaching set. The built-in examples include an untaught negative case so
the adapter is tested for both recall and over-generalization. LoRA files use a
format version; files produced before the
causal next-token label-alignment fix are rejected and must be retrained.
The strict evaluation separates these signals:
| Signal | Meaning |
|---|---|
after_grounded |
Verified-memory answer supplied by the Box controller |
after_raw_model |
The underlying model answered without fallback |
delayed |
Recall after simulated time and a paraphrased question |
restarted |
Recall after serializing and reloading memory |
memory_disabled |
Control for external-memory dependence |
negative_control |
The model did not claim the taught fact for an untaught entity |
repetition_ratio |
Fraction of repeated word tokens; lower is better |
Only a correct after_raw_model or post-LoRA memory_disabled result is
evidence that the model itself learned the fact. A grounded result alone is
useful product behavior, but it is retrieval, not parameter learning.
When evaluating a LoRA adapter, the script now computes
baseline_before_lora with a freshly loaded checkpoint before attaching the
adapter. This ordering is required: a baseline generated after loading LoRA
cannot measure learning because it is already adapted. The evaluator also
checks a paraphrased question with external memory disabled
(memory_disabled_paraphrase) to probe transfer beyond the exact training
prompt.
Run the corrected d4 experiment with:
python scripts/test_d4_checkpoints_in_box.py \
--device cuda \
--dtype float16 \
--max-new-tokens 64 \
--base-lora "./box_runs/online_d4_lora_v2.pt"The strongest result is a non-target baseline followed by correct
after_raw_model and memory_disabled answers, while the untaught negative
control remains incorrect. Even that result demonstrates only narrow
parameter adaptation on a synthetic fact; it is not evidence of general
knowledge acquisition or frontier-model-level intelligence.
After the single-fact smoke test, run a curriculum with six synthetic facts. The model trains on one canonical question per fact and is evaluated on different paraphrases with external memory disabled. Unknown entities are also tested to measure false-positive behavior.
python scripts/evaluate_d4_multifact.py \
--checkpoint "./d4/model_005140.pt" \
--device cuda \
--dtype float16 \
--steps 360 \
--rank 8 \
--positive-repeat 3 \
--negative-repeat 1 \
--seed 1234The report is written to box_runs/multifact_d4_report.json. Interpret the
metrics as follows:
train_accuracy: memorization of the taught canonical questions;heldout_accuracy: transfer to unseen paraphrases;negative_control_accuracy: abstention on entities never taught;mean_heldout_repetition: output quality indicator, where lower is better.safe_controller: a separate Box-level control using verified semantic memory to answer known facts and refuse unknown entities.
A useful learning signal is high train accuracy together with high held-out accuracy and negative-control accuracy. High train accuracy but low held-out accuracy indicates prompt memorization rather than robust learning. This experiment includes multiple question variants per fact and separate abstention examples so that unknown entities can be tested independently. A high train score with low held-out or negative-control scores indicates entity/color binding errors or over-generalization. This experiment is still synthetic; the next research step is to replace the facts with auditable public lessons and hold out concepts, not merely wording.
The first six-fact run (rank 4, 120 steps) reached train_accuracy=1.0,
heldout_accuracy=0.667, and negative_control_accuracy=0.0. The model learned
the canonical strings but swapped colors for two paraphrases and assigned
known colors to all unknown entities. The follow-up configuration above uses
more question variants, explicit abstention examples, rank 8, and a longer
training schedule to address those two failure modes.
The next rank-8 run reached train_accuracy=1.0 and
heldout_accuracy=1.0, confirming that the additional variants fixed the
entity/color binding errors. Its negative_control_accuracy remained 0.0:
all three unknown entities were assigned known colors. The curriculum now uses
12 abstention-training entities so the next run can measure whether explicit
open-world refusal reduces this hallucination without lowering held-out recall.
Because refusal answers are shorter and easier to optimize, the evaluator
oversamples positive examples by default (--positive-repeat 3 versus
--negative-repeat 1). A balanced run previously collapsed to universal
refusal (train_accuracy=0.167, heldout_accuracy=0.0,
negative_control_accuracy=1.0), so these sampling controls are explicit
experiment parameters rather than hidden heuristics.
Always record --seed when comparing runs. LoRA matrices are randomly
initialized; without a fixed seed, repeated commands can produce materially
different learning and hallucination rates.
To run the same experiment across several seeds and compute mean/stddev:
python scripts/evaluate_d4_multiseed.py \
--checkpoint "./d4/model_005140.pt" \
--device cuda \
--dtype float16 \
--steps 360 \
--rank 8 \
--positive-repeat 3 \
--negative-repeat 1 \
--seeds 1234,2024,42The aggregate file is box_runs/multiseed_d4/summary.json. Report the
parameter-learning metrics and the safe_controller metrics separately.
The first three-seed run produced:
| Metric | Mean | Std. dev. |
|---|---|---|
| Train accuracy | 0.9444 | 0.0786 |
| Held-out accuracy | 0.9444 | 0.0786 |
| Pure-model negative control | 0.0000 | 0.0000 |
| Safe-controller positive accuracy | 1.0000 | 0.0000 |
| Safe-controller negative accuracy | 1.0000 | 0.0000 |
This is strong evidence of narrow positive-fact adaptation, but also a clear failure of model-level open-world refusal. The Box therefore treats verified memory gating as a separate safety layer instead of claiming that LoRA alone solves hallucination.
The Box can ingest auditable lesson records without using another language
model as the teacher. Each JSONL record must include a lesson ID, an original
summary, a source URL, and a license. Optional concepts and prerequisites are
validated before teaching. The sample file contains eight short original summaries
linked to MIT OpenCourseWare pages; it is not a bulk copy of course material.
Curriculum titles and concepts are stored as provenance tags so safe retrieval
can match a question such as “What is Newton's second law?” to the correct
verified fact instead of relying on generic words like force or color.
Teach the sample curriculum into a durable SQLite memory database:
python scripts/teach_public_curriculum.py \
--curriculum "./curricula/mit_ocw_intro_sample.jsonl" \
--checkpoint "./d4/model_005140.pt" \
--device cuda \
--dtype float16 \
--verify \
--ask "What is Newton's second law?"The script writes box_runs/public_curriculum.sqlite3 and an append-only event
log. --verify is an explicit operator decision that the cited source was
reviewed; importing a lesson alone never marks it as verified. Re-running the
command skips semantic facts with the same source citation.
Run a post-restart curriculum exam against the imported SQLite database:
python scripts/evaluate_public_curriculum.py \
--curriculum "./curricula/mit_ocw_intro_sample.jsonl" \
--checkpoint "./d4/model_005140.pt" \
--device cuda \
--dtype float16The exam reports lesson accuracy, unknown-entity refusal, restart accuracy, and whether every answer retrieved a memory with the expected source URL. A CPU smoke exam reached 1.0 for all three accuracies and matched all sources.
To run the curriculum as a staged life, use the prerequisite-aware progress runner. It opens at most one lesson per invocation by default, administers the lesson exam, and records completion only after a passing answer:
python scripts/run_curriculum_life.py \
--curriculum "./curricula/mit_ocw_intro_sample.jsonl" \
--checkpoint "./d4/model_005140.pt" \
--device cuda \
--dtype float16 \
--verify \
--max-lessons 1Progress is stored in box_runs/curriculum_progress.json, while knowledge and
provenance remain in box_runs/public_curriculum.sqlite3. Run the command
again to advance to the next ready lesson; failed exams remain incomplete and
can be retaken.
The sample includes intermediate lessons for kinematics, the calculus chain
rule, consistent linear systems, and gene expression; each is unlocked only
after its corresponding introductory lesson is completed.
The report also includes safe_controller. This is intentionally not counted
as parameter learning: it measures whether the Box can enforce reliable
behavior using provenance-aware verified memory even when the underlying model
would hallucinate.
After all curriculum exams pass, consolidate the public lessons into a d4 LoRA adapter with a separate memory-disabled evaluation:
python scripts/consolidate_public_curriculum_d4.py \
--curriculum "./curricula/mit_ocw_intro_sample.jsonl" \
--progress "./box_runs/curriculum_progress.json" \
--checkpoint "./d4/model_005140.pt" \
--memory-db "./box_runs/public_curriculum.sqlite3" \
--device cuda \
--dtype float16 \
--steps 480 \
--rank 8 \
--positive-repeat 3 \
--negative-repeat 0 \
--seed 1234The script refuses to train if any curriculum exam is incomplete. Its
adapted_accuracy measures parameter recall with external memory disabled;
reloaded_accuracy repeats that measurement after loading the saved LoRA into
a fresh checkpoint instance;
safe_accuracy measures the separate verified-memory controller. These must
remain separate when reporting results. Consolidation trains on both each
lesson's exam question and a title-based explanation prompt. Refusal examples
are optional (--negative-repeat 0 by default) because they previously caused
the adapter to collapse into universal refusal; unknown-entity safety is handled
by the verified-memory controller and evaluated independently.
The first public consolidation run reached adapted_accuracy=1.0 and
safe_accuracy=1.0 with negative_repeat=0. The zero pure-model negative score
is expected for this positive-only consolidation objective; it must not be
interpreted as a failure of the Safe Controller. The evaluator now also emits
reloaded_accuracy, which is the required persistence check for the saved
adapter.
To test continual learning and catastrophic forgetting, run two sequential phases: first consolidate the eight completed public lessons, then train on four unrelated facts without replaying the old examples:
python scripts/evaluate_d4_retention_forgetting.py \
--curriculum "./curricula/mit_ocw_intro_sample.jsonl" \
--progress "./box_runs/curriculum_progress.json" \
--checkpoint "./d4/model_005140.pt" \
--device cuda \
--dtype float16 \
--phase1-steps 480 \
--phase2-steps 240 \
--rank 8 \
--positive-repeat 3 \
--seed 1234The report compares old-knowledge accuracy after phase 1, after phase 2, and
after reloading the final adapter, together with accuracy on the new facts.
old_knowledge_retained is true only when phase 1 is already meaningful and
the reloaded score does not drop by more than 25 percentage points.
To test experience replay, rerun the same command with
--replay-repeat 3. Phase 2 then mixes old examples with new facts; compare
phase2_old_accuracy and phase2_new_accuracy against the no-replay report.
The phase-2 trainer now builds an exactly phase2_steps-long deterministic
shuffled schedule, so a non-divisible step count cannot bias the final updates
toward only the newest examples.
The first replay run achieved phase1_old_accuracy=1.0,
phase2_old_accuracy=0.875, phase2_new_accuracy=1.0, and identical scores
after reloading the final adapter. Only the chain-rule item was forgotten,
compared with 87.5% forgetting in the no-replay run. This is initial evidence
that replay substantially reduces interference, not proof that forgetting is
solved.
After the deterministic phase-2 schedule fix, replay achieved 1.0 for old knowledge, new knowledge, and both reloaded scores in the same 8-old/4-new experiment. This confirms that the schedule correction removed the previous chain-rule loss for this seed. Multi-seed retention and a larger curriculum are still required before claiming robust continual learning.
For the next stability check, evaluate_d4_retention_multiseed.py repeats the
two-phase experiment across several seeds and replay ratios. The default
configuration includes a no-replay baseline (0) and replay ratios 1, 3,
and 5:
python scripts/evaluate_d4_retention_multiseed.py \
--curriculum "./curricula/mit_ocw_intro_sample.jsonl" \
--progress "./box_runs/curriculum_progress.json" \
--checkpoint "./d4/model_005140.pt" \
--device cuda \
--dtype float16 \
--phase1-steps 480 \
--phase2-steps 240 \
--rank 8 \
--positive-repeat 3 \
--replay-ratios 0,1,3,5 \
--seeds 1234,2024,42Runs are sequential to fit a single GPU. Each run writes its own JSON report
and LoRA files; box_runs/retention_multiseed/summary.json reports the
mean, population standard deviation, and per-seed values for old/new accuracy
before and after reload. Prefer the smallest replay ratio whose reloaded old
accuracy remains high without reducing new-fact accuracy. A single perfect
seed is encouraging, but a ratio should only be considered stable when its
mean is high and its standard deviation is small across all seeds.
The first three-seed run (seeds 1234,2024,42) produced this result:
| Replay repeat | Reloaded old accuracy | Reloaded new accuracy |
|---|---|---|
| 0 | 0.0000 | 1.0000 |
| 1 | 1.0000 | 1.0000 |
| 3 | 1.0000 | 1.0000 |
| 5 | 1.0000 | 1.0000 |
All values had zero standard deviation across the three seeds in this small experiment. This is strong evidence that replay is necessary for retention in this setup, but the task is still only twelve short factual items. The next evaluation should test conflicting or corrected facts, broader concept coverage, and held-out reasoning rather than treating this result as evidence of general intelligence.
Lifelong learning must handle knowledge revisions, not only append new facts.
SAssistantBox.verify_fact accepts an optional supersedes memory ID. The new
fact is verified, the old fact remains available for audit, and the old fact is
marked superseded and excluded from safe/grounded answers. When independently
verified facts still conflict, the controller chooses the highest-confidence
fact deterministically; an exact-confidence tie chooses the later retrieved
item. Every verification and supersession is written to the event log.
Run the deterministic controller test with:
python scripts/evaluate_box_conflicts.pyThe report is written to box_runs/conflicts/report.json. It checks that a
correction wins over the old answer, confidence resolves an unlinked conflict,
unknown entities are rejected, the correction survives a memory restart, and a
fact_superseded event is recorded. This test evaluates Box provenance and
retrieval policy, not d4 parameter learning, so it must be reported separately
from LoRA accuracy.
The initial conflict run passed all seven checks. The corrected answer was
crimson instead of the old violet, the independent confidence conflict
selected blue over amber, the unknown entity returned I do not know.,
and the corrected result remained available after restarting from JSON memory.
evaluate_public_transfer.py evaluates the verified memory using held-out
question wording rather than the exact curriculum exam questions. It also
reopens the SQLite database and repeats the questions after restart. The
entity-aware safety controller now requires a distinctive subject match;
generic words such as derivative, information, or matrix cannot select a
fact by themselves. Explicit unknown markers are rejected as well.
python scripts/evaluate_public_transfer.py \
--curriculum "./curricula/mit_ocw_intro_sample.jsonl" \
--checkpoint "./d4/model_005140.pt" \
--device cuda \
--dtype float16 \
--memory-db "./box_runs/public_curriculum.sqlite3" \
--output "./box_runs/public_transfer.json"The first completed run achieved 1.0 held-out transfer accuracy, 1.0
restart accuracy, 1.0 unknown-question accuracy, grounded answers for all
held-out items, and matching public sources. The run used CPU only because the
validation environment had no CUDA device; this does not represent GPU
performance. The final validation report is box_runs/public_transfer_v5.json.
The next curriculum is curricula/mit_ocw_expanded_sample.jsonl. It contains
16 original MIT OCW-linked summaries across mechanics, calculus, linear
algebra, and biology, with prerequisite edges reaching advanced topics. The
separate curricula/mit_ocw_expanded_transfer.jsonl file contains 16 held-out
concept-transfer questions and must not be imported as teaching material.
python scripts/teach_public_curriculum.py \
--curriculum "./curricula/mit_ocw_expanded_sample.jsonl" \
--checkpoint "./d4/model_005140.pt" \
--device cuda \
--dtype float16 \
--verify \
--memory-db "./box_runs/public_curriculum_expanded.sqlite3" \
--event-log "./box_runs/public_curriculum_expanded.jsonl"
python scripts/evaluate_public_transfer.py \
--curriculum "./curricula/mit_ocw_expanded_sample.jsonl" \
--questions "./curricula/mit_ocw_expanded_transfer.jsonl" \
--checkpoint "./d4/model_005140.pt" \
--device cuda \
--dtype float16 \
--memory-db "./box_runs/public_curriculum_expanded.sqlite3" \
--output "./box_runs/public_transfer_expanded.json"The first expanded run taught all 16 lessons and stored 16 verified semantic
facts. After retrieval normalization was strengthened, CPU validation reached
1.0 held-out transfer accuracy, 1.0 restart accuracy, 1.0 unknown
accuracy, grounded answers for all items, and matching sources. This remains a
small benchmark; the next curriculum must use longer source-backed lessons and
genuinely novel problem-solving tasks.
The Box now has a small, explicit reasoning layer in src/s_assistant_box/reasoning.py.
It provides auditable procedures for force (F = ma), constant-acceleration
velocity, power derivatives and integrals, matrix-product dimensions, momentum
conservation, and mechanical-energy conservation. A procedure runs only when
the required lesson ID is present as a verified semantic memory. Missing or
invalid prerequisites return I do not know. rather than silently inventing
an answer.
Run the end-to-end check against the expanded SQLite curriculum:
python scripts/evaluate_box_reasoning.py \
--memory-db "./box_runs/public_curriculum_expanded.sqlite3" \
--output "./box_runs/reasoning/report.json"The initial run passed all seven exercises and the missing-prerequisite refusal check. This is explicitly tool-assisted reasoning, not evidence that d4 learned arithmetic in its parameters. Future experiments should compare this layer against pure d4 generation, retrieval-only answers, and LoRA-consolidated answers under equal question and compute budgets.
evaluate_box_capability_sources.py runs the same seven bounded exercises in
three modes: d4 generation with no memory, verified-memory safe answers, and
explicit procedures gated by verified lesson IDs. This identifies whether an
answer came from model parameters, Box retrieval, or a tool.
python scripts/evaluate_box_capability_sources.py \
--checkpoint "./d4/model_005140.pt" \
--memory-db "./box_runs/public_curriculum_expanded.sqlite3" \
--device cuda \
--dtype float16 \
--output "./box_runs/capability_sources.json"The first CPU validation produced model-only accuracy 0.0, verified-memory
accuracy 0.2857, and tool-assisted accuracy 1.0. All tool procedures were
accepted after verified-prerequisite checks. This exposes a useful limitation:
lexical retrieval is not sufficient for numerical/compositional questions.
The perfect tool score measures the explicit procedure layer, not d4
intelligence, and must be reported separately.
The reasoning path is now integrated into SAssistantBox.reason(task). Each
call checks the durable memory, runs the procedure, and appends a reasoning
event containing the task inputs, result, acceptance flag, and evidence memory
IDs. The capability-source report was rerun through this path with unchanged
scores (0.0 model-only, 0.2857 verified-memory, 1.0 tool-assisted), and
the full test suite passes 29 tests.
Reasoning results can now enter a durable positive replay buffer only after an
external evaluator (or user) marks them correct. SAssistantBox.reason()
produces the auditable result; submit_reasoning_feedback() records the
decision as a reasoning_feedback event and adds accepted examples with their
evidence memory IDs. Incorrect or unaccepted results remain in the event log
for analysis but are never used as positive replay data.
Run the feedback-loop check:
python scripts/evaluate_reasoning_feedback.py \
--memory-db "./box_runs/public_curriculum_expanded.sqlite3" \
--replay "./box_runs/reasoning/replay.json" \
--output "./box_runs/reasoning/feedback_report.json"The initial run passed all five checks: correct feedback was added, incorrect
feedback was rejected, replay survived JSON reload, both feedback events were
logged, and Box status reported the replay count. The full suite now passes 29
tests. This buffer is ready to feed a future LoRA/replay scheduler, but it is
now connected by train_d4_from_feedback.py to an explicit, reproducible d4
LoRA update:
python scripts/train_d4_from_feedback.py \
--checkpoint "./d4/model_005140.pt" \
--replay "./box_runs/reasoning/replay.json" \
--device cuda \
--dtype float16 \
--steps 120 \
--rank 8 \
--replay-repeat 3 \
--output "./box_runs/reasoning/feedback_lora.pt" \
--report "./box_runs/reasoning/feedback_lora_report.json"The trainer measures memory-disabled d4 accuracy before and after the update, stores the final loss and seed, and preserves each replay example's task ID source, and numeric inputs. Including inputs is essential for calculation replay: a model must see the same problem instance that produced the accepted answer. This is parameter adaptation from evaluator-approved feedback; it does not turn tool-generated answers into unverified model knowledge.
Every feedback prompt includes an explicit Procedure: field. This prevents
short questions such as Find force and Find velocity from being treated as
the same task and makes cross-procedure interference measurable.
Do not interpret a very small final loss as proof that the model learned the procedure. Autoregressive training can overfit one completion while free generation still repeats a token sequence. Evaluate the complete replay and separate teacher-forced metrics from generation metrics:
python scripts/train_d4_from_feedback.py \
--checkpoint "./d4/model_005140.pt" \
--replay "./box_runs/reasoning/replay_full.json" \
--device cuda \
--dtype float16 \
--steps 120 \
--rank 8 \
--replay-repeat 3 \
--output "./box_runs/reasoning/feedback_lora_full.pt" \
--report "./box_runs/reasoning/feedback_lora_full_report.json"Then run the strict, memory-disabled evaluator:
python scripts/evaluate_d4_feedback.py \
--checkpoint "./d4/model_005140.pt" \
--replay "./box_runs/reasoning/replay_full.json" \
--lora "./box_runs/reasoning/feedback_lora_full.pt" \
--device cuda \
--dtype float16 \
--max-new-tokens 48 \
--output "./box_runs/reasoning/feedback_eval_full.json"The report contains exact replay accuracy, paraphrase accuracy, held-out numeric-input accuracy, teacher-forced answer-token accuracy/loss, and a fresh LoRA reload check. A successful feedback experiment should improve exact and held-out free generation, retain a high teacher-forced score, and reproduce the result after reload. If only teacher-forced loss improves, classify the run as memorization/overfitting rather than learned reasoning.
For numeric procedures, expand the accepted replay before training. The
additional rows are deterministic examples generated by the registered
procedure (including negative boolean and incompatible-matrix cases), and are
marked as synthetic in their source field:
python scripts/augment_reasoning_replay.py \
--input "./box_runs/reasoning/replay_full.json" \
--output "./box_runs/reasoning/replay_augmented.json" \
--variants-per-procedure 4 \
--seed 1234The first full-replay run (7 accepted examples, 120 steps, rank 8) reached
0.2857 free-generation accuracy and 0.8867 teacher-forced token accuracy.
A 480-step run reached 0.8571 exact replay accuracy and 0.9911
teacher-forced accuracy, but only 0.2857 on unseen numeric inputs. Its
outputs included incorrect arithmetic such as F = 3 * 4 = 13 N.
An augmented run (31 rows with additional numeric and negative-branch cases,
960 steps) reached 0.4286 held-out accuracy and 0.9655 teacher-forced token
accuracy. Adding an explicit Procedure: field to every prompt did not remove
the remaining cross-procedure errors. These measurements show that a low SFT
loss can coexist with poor autoregressive calculation and should be reported
as overfitting/limited parameter adaptation, not as learned general reasoning.
For exact arithmetic and prerequisite-gated answers, the verified Box tools
remain the authoritative path; LoRA is an optional learned shortcut whose
quality must pass the strict evaluator.
The latest source ablation on the d4 Base checkpoint measured 0.0
model-only accuracy, 0.2857 verified-memory accuracy, and 1.0
tool-assisted accuracy. The tool path accepted all seven tasks; this is a
capability decomposition, not evidence that the d4 parameters learned the
procedures.
The Box now exposes SAssistantBox.reason_question(question). A deterministic
router recognizes the supported force, kinematics, calculus, matrix, momentum,
and mechanical-energy forms, extracts numeric or boolean inputs, checks the
corresponding verified prerequisite lesson, and invokes the existing procedure
tool. It never asks the language model to perform an unverified calculation.
Incomplete, ambiguous, or unsupported questions return I do not know. and
are logged as reasoning_route events.
Run the router evaluation against the expanded public curriculum database:
python scripts/evaluate_reasoning_router.py \
--memory-db "./box_runs/public_curriculum_expanded.sqlite3" \
--output "./box_runs/reasoning/router_report.json"The current evaluation routes all seven supported questions correctly, finds verified evidence for every answer, and refuses all ambiguous/unsupported controls. This establishes a safe natural-language interface around the Box tools while keeping d4 parameter learning as a separately measured experiment.
SAssistantBox.reason_chain() executes an ordered list of ReasoningTask
objects. A later task may reference a structured value from an earlier task
with the syntax $step_index.field, for example $0.final_velocity. Every
step performs its own prerequisite-memory check; an unresolved reference or a
missing lesson aborts the chain and returns I do not know.. The event log
stores the chain, individual answers, and the union of evidence IDs.
Example evaluation:
python scripts/evaluate_reasoning_chains.py \
--memory-db "./box_runs/public_curriculum_expanded.sqlite3" \
--output "./box_runs/reasoning/chains_report.json"The current chain test computes final velocity (13 m/s) and feeds it into a
kinetic-energy step (253.5 J); it also verifies safe refusal for an invalid
step reference. Both checks pass without using the language model.
d4 is the primary scientific experiment because it is a Base Model. d4-sft
is useful as a practical assistant baseline, but its instruction-following
behavior comes from supervised fine-tuning that happened before Box import.
Both supplied d4 checkpoints use n_layer=4, n_embd=256, and a 32,768-token
vocabulary (about 36.7M stored parameters), so they are not the 124M GPT-2-style
model trained by base model/train_gpt2.py. Keep the two model families as
separate experiments and do not compare their raw losses directly.
The first kernel is intentionally dependency-light. It can run with a callable model adapter before PyTorch and Transformers are installed.
python -m pip install -e ".[dev]"
python -m pytest
python -m s_assistant_box.cliTo connect a local Hugging Face model, install the model extra and create a
TransformersModelAdapter with a ModelConfig. The same code path supports
CPU and CUDA; select a device explicitly with HardwareConfig(device="cpu"),
HardwareConfig(device="cuda"), or leave it as "auto".
For the included 124M training script:
cd .
python -m pip install -e ".[training]"
cd "./base model"
python train_gpt2.py --device cuda --precision auto --micro-batch-size 0With --micro-batch-size 0, the script selects a conservative value from the
detected GPU memory. On an RTX 3070 it normally selects 4 samples at sequence
length 512 and uses FP16 with gradient scaling. On an H100 it can select a
larger micro-batch and prefer BF16. Reduce the micro-batch size or sequence
length if a particular driver or CUDA build reports an out-of-memory error.
For additional RTX 3070 throughput, test a larger micro-batch after the smoke test:
python train_gpt2.py --device cuda --precision fp16 --micro-batch-size 8 --sequence-length 512 --max-steps 100 --total-batch-size 65536 --hellaswag-interval 0 --sample-interval 0Keep the setting only if the run remains stable and GPU memory stays below the
limit. --compile is available as an optional optimization, but its startup
overhead and Windows support should be benchmarked on the local PyTorch build.
Useful smoke-test options are:
python train_gpt2.py --device cuda --max-steps 100 --total-batch-size 65536 --hellaswag-interval 0 --sample-interval 0Checkpoints include model weights, optimizer state, mixed-precision scaler
state, random-number state, and data-loader position. The filename represents
the number of completed optimizer steps. Resume with, for example,
--resume path/to/model_000100.pt.
- define the model interface;
- define the first artificial world;
- define teaching and evaluation data;
- define the compute budget;
- define success criteria.
- load the base model;
- implement the interaction loop;
- implement simulated time and events;
- persist all events and model outputs.
- implement working memory;
- implement episodic memory;
- implement semantic memory;
- implement procedural memory;
- add retrieval, confidence, decay, and provenance.
- implement parent and teacher roles;
- implement prerequisite-aware curriculum;
- add exercises, feedback, review, and exams;
- add stage transitions.
- summarize episodes;
- merge and verify concepts;
- implement spaced replay;
- add optional adapter or LoRA training;
- test forgetting and relearning.
- run controlled baseline comparisons;
- measure retention and transfer;
- profile CPU, memory, and storage requirements;
- publish reproducible experiment configurations.
- add interests and career selection;
- add projects and research;
- add social relationships;
- add long-term goals and reflection;
- add a visual interface.
The general direction has important precedents. Related areas include:
- continual learning and lifelong learning;
- curriculum learning;
- reinforcement learning and learning through interaction;
- retrieval-augmented generation;
- neural episodic memory;
- differentiable memory systems;
- cognitive architectures such as ACT-R and Soar;
- generative agents with memory and reflection;
- simulated environments and embodied agents;
- autonomous research and self-improvement systems.
Therefore, the broad idea of giving an agent experience, memory, goals, and staged learning has been explored by other researchers.
The potentially distinctive contribution of S Assistant Box is the specific combination and experimental framing:
- placing a very small language model inside an explicit life-course environment;
- separating parental, educational, professional, and research roles;
- treating memory consolidation as a central learning mechanism;
- measuring learning under a controlled compute budget;
- testing whether structured life experience is more efficient than one-shot static training.
The novelty should be established through implementation details, ablation studies, and reproducible results rather than through the life metaphor alone.
The repository is designed to publish source code and small synthetic curriculum examples, not private experiment state. Before creating a public repository:
- review
git diff --cachedandgit ls-filesfor unexpected files; - keep checkpoints, optimizer states, downloaded datasets, SQLite databases, event logs, and generated run artifacts out of Git;
- never commit API keys, access tokens, passwords, private URLs, or personal screenshots;
- pass local model paths with
--checkpointand--nanochat-rootinstead of hard-coding them in source code; - verify the license and redistribution terms of every external dataset, checkpoint, tokenizer, and source excerpt;
- use synthetic data for tests and issue reproductions;
- enable GitHub secret scanning and review its alerts after the first push.
The included .gitignore excludes model files, processed datasets, local
box_runs/, databases, caches, and environment files. It intentionally keeps
the small curriculum JSONL files because they contain original summaries and
links rather than mirrored course content.
This repository contains the project specification and a runnable Box Kernel.
The kernel provides life-stage state, append-only event logging, portable device
detection, GPT-2 and nanochat d4 model adapters, episodic and semantic memory,
fact provenance and verification, JSON restart snapshots, hybrid and safe
grounded answers, strict memory-disabled evaluation, and online LoRA SFT for d4.
It also includes a durable SQLiteMemoryStore with the same memory API as the
in-memory/JSON store, so long-running Box sessions can survive process restarts
without loading the entire memory into Python objects.
The safe answer mode returns I do not know. when no relevant fact has been
verified, preventing a generative model from assigning an unrelated known fact
to an unknown entity. Relevance is entity-aware: a verified fact must share a
specific subject token with the question, not merely generic words such as
mineral or color. This controller-level safety behavior is evaluated
separately from pure parameter-learning tests. The next milestones are larger
held-out curricula, public-source ingestion, and longitudinal
adapter/consolidation experiments. A multi-seed retention/replay runner is now
available; its aggregate report is required before drawing broad continual-
learning conclusions.
The feedback path now also includes strict generation-vs-teacher-forcing
evaluation (scripts/evaluate_d4_feedback.py) and deterministic procedure
replay augmentation (scripts/augment_reasoning_replay.py).
This project is released under the MIT License. See LICENSE.
External datasets, checkpoints, tokenizers, course materials, and source summaries may have separate licenses. Check their original terms before redistributing them. See NOTICE for the third-party attribution summary.