Back to AI guides

mini-AGI: A Byte-Level Language Model That Trains Continuously on 8GB VRAM

Stanley Ulili
Updated on October 5, 2026

Most language models stop learning once training is complete. Their knowledge is largely limited to what was captured during training, and adding substantial new knowledge typically means another training run. At frontier scale, that can require weeks of GPU time and cost tens of millions of dollars. Fine-tuning is cheaper, but it introduces another problem: catastrophic forgetting, where learning new information can degrade knowledge the model already had.

mini-AGI, created by Alexey Borsky, explores a different approach. It is a byte-level language model designed to learn continuously from a stream of data. The model reads a chunk, takes a learning step, moves to the next chunk, and keeps training rather than becoming fixed after an initial training phase.

The goal is to let the model add new knowledge without erasing what it learned before, while keeping the hardware requirements relatively modest. mini-AGI is designed to run on a single consumer GPU with 8 GB of VRAM. Despite the ambitious idea behind the project, Borsky describes the name "mini-AGI" as "a half-joke."

What makes it different

Byte-level, no tokenizer. The alphabet is 256 byte values. There's no tokenizer to fit, no vocabulary to predefine, and no data type that needs special handling. Any text, code, or byte sequence can be fed in directly.

Batch size 1, continuous stream. Standard training collects data into large batches and shuffles them before each pass. mini-AGI reads one chunk at a time in sequence, takes one gradient step, and moves to the next chunk. The learning rate is not scheduled with a cosine decay that assumes training will end. It's governed dynamically by plasticity.py based on held-out loss, because for a model designed to train indefinitely, a schedule that implies an endpoint doesn't make sense.

Disk-paged Mixture of Experts. The model stores individual expert networks as files on disk. When processing a chunk of text, the model loads only the relevant experts into VRAM, runs them, and pages them back out. Total disk storage can far exceed VRAM capacity. The current live run has 174 experts across multiple domains.

Dynamic architecture. The model grows and prunes itself as it trains. When genuinely new material appears that existing experts handle poorly, the model spawns a new expert. When experts go unused, they get pruned to reclaim disk space. The architecture you end up with is not the one you started with.

The official GitHub repository page for mini-AGI, showing the project's description and core philosophy.

How the catastrophic forgetting problem is addressed

The core claim of mini-AGI is that it can learn new domains without erasing old ones. The architecture uses two learning rates:

The trunk is the shared core that every input passes through. It learns fundamental patterns, grammar and syntax. It trains at one-tenth the learning rate of the experts.

The experts are the specialist sub-networks for specific topics and styles. They train at the full learning rate and adapt quickly to new material.

When the model learns a new domain, the experts for that domain update rapidly while the trunk changes very slowly. When it returns to old material, the relevant experts reactivate, and because the trunk still carries the foundational patterns, performance recovers quickly.

The README's own measured cost: "0.13 nats across eight domains to gain 0.42 on a ninth." Learning genuinely new material comes at a real but small cost to existing knowledge, and the gain substantially outweighs it.

One honest caveat from a community replication: a user found that running all layers at 0.1x (no trunk/expert split) produced similar forgetting rates to the split approach over a 4M character run, and that chess knowledge grew only about 0.04 over that stretch. The author's response was that the trunk/expert split becomes more important over longer runs and with more domain switches. The project's README is, by its own description, "unusually candid about what was tried and what failed," and this kind of open community scrutiny is part of how the project is developing.

The two-phase experiment

The Tarantino experiment in the source article illustrates the system clearly, so it's worth covering concretely.

A diagram illustrating the two-step training plan, starting with the "TinyStories" dataset to learn English, then moving to Tarantino's screenplays.

Phase 1: Start from completely random weights and train on TinyStories, a collection of simple children's stories generated by GPT-3.5 and GPT-4 with small vocabulary and simple sentence structure. Within a few minutes the model produces gibberish. After a few hours it produces coherent sentences.

Phase 2: Switch the training stream to Tarantino screenplays (Pulp Fiction, Jackie Brown, Inglourious Basterds, Django Unchained, Kill Bill, The Hateful Eight, Four Rooms). Validation uses a held-out Reservoir Dogs screenplay the model has never seen.

Within five minutes of seeing the screenplay format, the model adopts it, with character names in caps and indented dialogue. The styles clash for a while in ways the README describes honestly: characters from the Tarantino universe produce polite kindergartener sentences, phrases like "Stephen, I'm sorring."

The validation graph is the key result:

A line graph showing the "English score" over time. The score improves, worsens during the "Tarantino" phase, and then quickly recovers and improves further when switched back to stories.

English performance improves during Phase 1, degrades when the training switches to Tarantino, then recovers to its previous level within seven minutes of returning to children's stories and continues improving past its previous best. The foundational knowledge was suppressed but not erased.

Setup and training

Requirements: Python, PyTorch, and an 8GB VRAM GPU (or CPU, much slower).

Clone the repository:

 
git clone https://github.com/volotat/mini-AGI.git
cd mini-AGI

Install dependencies in a virtual environment:

 
pip install -r requirements.txt

Prepare text files for your training data. For the two-phase experiment: tinystories.txt for the English phase, tarantino.txt for the Tarantino phase, and heldout.txt (Reservoir Dogs) for validation. Configure the training phases and dataset paths in config.yaml, then start:

 
python train.py

The training script starts a local dashboard at http://127.0.0.1:8788.

The mini-AGI training control dashboard, displaying key metrics like characters read, held-out loss, stories and tarantino scores, and reading speed.

The dashboard shows characters read, held-out loss (lower is better), domain-specific scores, and reading speed. It also shows live samples of what the model is currently generating, so you can watch it improve in real time.

The current reference run (not the Tarantino experiment but the author's main ongoing run) is at 828.7M characters over 1,503 evaluations, with a held-out loss of 0.6735 nats. Pre-trained weights are available on Hugging Face as Volotat/mini-AGI-undertrained for experimentation without running your own training from scratch.

The interactive interface

After training, mini-AGI includes a screenplay-format chat interface.

The screenplay chat web application, an interface for interacting with the trained mini-AGI model in a conversational, script-writing format.

The interface is framed as a script collaboration: you write a line, the model continues in character. After Tarantino training the outputs are imperfect and often funny, but the stylistic patterns are recognizably there: the cadence, the profanity, the extended dialogue riffs.

What this is and isn't

mini-AGI is a proof-of-concept research project, not an attempt to compete with commercial LLMs. The model is deliberately small, and its writing quality after training reflects that scale. The architecture is built to explore continual learning rather than maximize benchmark scores or generation quality.

What matters is the learning behavior it demonstrates. A language model can continuously learn from a single stream of data on one GPU with 8 GB of VRAM, pick up new domains without completely erasing what it learned before, and adjust its own architecture by growing or shrinking as the data changes.

That combination is the interesting part. Continuous learning, resistance to catastrophic forgetting, and dynamic model growth are brought together in a system small enough to experiment with on consumer hardware, rather than requiring the infrastructure behind a frontier-scale training run.

Got an article suggestion? Let us know
Licensed under CC-BY-NC-SA

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.