M1-128M
From-scratch hybrid — attention + liquid core in every layer. 50K-step run on WikiText-103. English text continuation.
Best for English text continuation & inference-cost demos.
Every released model, with its exact parameters and download link. Run them on your own hardware — CPU-friendly, no GPU required.
From-scratch hybrid — attention + liquid core in every layer. 50K-step run on WikiText-103. English text continuation.
Best for English text continuation & inference-cost demos.
Frozen TinyLlama-1.1B-Chat base with trained liquid-core residual adapters + LoRA. The conversational model — English and Chinese.
Best for Bilingual conversation with cross-turn memory.
Attention-free O-series edge model. Constant 0.381MB carried state regardless of context — the O(1) memory line. Runs on CPU.
Best for Always-on edge streaming, CPU-only inference.
1.93B-parameter M2 preset (2080d × 34L × 16 heads), byte-level tokenizer. The baseline run completed at 200,000 steps — final val PPL 2.4106, context-flat (2.72 / 2.72 / 2.68 at 128 / 256 / 512). A language-scaling research milestone, not an instruct model.
Best for Byte-level language-scaling research.
MT-LNN residual adapter on a frozen Qwen2.5-0.5B-Instruct, mounted at 6 of 24 layers, with 12.4M trainable parameters (2.5%). Initialised near-identity so the base capability is preserved; a deliberately short 500-step research checkpoint, not a converged model.
Best for Research on liquid adapters over an instruct base.
Always-on wake-word spotter on the O-Series liquid core. Holds a constant 5,120-byte carried state no matter how long the microphone stays open, and fires on a greeting ("hello", "hola", "bonjour", "你好"). Streaming ONNX graph: 5.03 MB fp32 / 1.27 MB int8.
Best for Battery-friendly always-on keyword spotting.
Sparse spiking neural networks whose structural masks follow the statistical laws of the real Drosophila connectome (5% density, LIF neurons, surrogate-gradient training). MNIST 96.83%, Fashion-MNIST 87.07% — near-dense accuracy at an estimated ~107–112× lower energy.
Best for Energy-efficient edge classification research.
One glance: what each model is for, how big it is, and where to get it.
| Model | Best for | Params | Window | Carried state | Status |
|---|---|---|---|---|---|
| M1-128M | English text continuation | 128.6M | 512 tokens | 0.381 MB (O(1)) | ● Live on site |
| M1 · TinyLlama adapter | Bilingual conversation, cross-turn memory | 1.1B + 2.3M | — | — | ● Live on site |
| O1-48M | Always-on edge streaming, CPU-only | 48M | — | 0.381 MB (O(1)) | ↓ Download |
| M2-2B | Byte-level language-scaling research | 1.93B | 512 tokens | — | ↓ Download |
| O1-Qwen05-Adapter | Liquid adapter on an instruct base (research) | 0.5B + 12.4M | — | — | ↓ Download |
| O1-Sound | Always-on wake-word spotting | ~1.3M | — | 5,120 B (O(1)) | ↓ Download |
| Sparse-SNN | Energy-efficient edge classification | ~1.6M | — | — | ↓ Download |
Weights on HuggingFace (AwareLiquid). More research lines (O1-Anti, C1, AwareLiquid-Physic, human-brain-simulation): clone & run from github.com/AwareLiquid.
git clone https://github.com/AwareLiquid/M1 && cd M1
pip install -r requirements.txt
# drop the downloaded .pt into checkpoints/, then:
CKPT_PATH=checkpoints/hybrid_125m_serve.pt TOKENIZER=gpt2 \
python -m uvicorn serve.server:app --port 8000
All models run CPU-only. The 128M hybrid needs ~1GB RAM; the TinyLlama adapter needs ~5GB. Full instructions in the repo README.