New · Terse Micro · Available now

423M params.
182 MB.
No GPU.

Terse Micro is a ternary-weight language model built from scratch — every weight is −1, 0 or +1. Trained end-to-end on 8 billion tokens, it ships as a single ~182 MB file and runs on a CPU — no GPU, no cloud. Apache 2.0.

⚑ Trained end-to-end (8B tokens · SFT · ORPO) · single-file GGUF on Hugging Face · Apache-2.0

terse-micro · TQ2_0
# the whole model — one file
terse-micro.TQ2_0.gguf   ~182 MB

architecture   MoE · 4 experts, top-2
parameters     ~423M  (~320M active/token)
weights        ternary  {−1, 0, +1}
runtime        CPU  no GPU needed
training       8B tokens  from scratch
license        Apache-2.0
−10+1 three values · real intelligence

423M params
Mixture-of-experts — ~320M active per token
182MB file
The entire model, TQ2_0-quantized into one file
1.6bits/wt
Ternary weights — non-embedding weights ~10× smaller than fp16
8B tokens
Trained from scratch on FineWeb-Edu (~19 tok/param)

The Terse family

Ternary models, built
from scratch.

See the full family
Terse Micro
Available now

~423M params (~320M active). Runs on any phone or CPU. ~182 MB file.

text · phone / CPU
Terse Mini Planned

Laptop-class — the next step up the ladder.

text · laptop
Terse Medium-Lite Planned

Runs on an ordinary 16 GB-RAM laptop — no discrete GPU.

text · 16 GB laptop
Terse Medium Planned

Multimodal, on a gaming laptop — 16 GB RAM + 4 GB GPU.

text + image · gaming laptop
Terse Pro Planned

Frontier-scale on a single server node. Fully multimodal.

text + image + video · server

Also from Ternative

Orchid 1.0 ranks
#3 — behind only
Claude and GPT-4o.

Our earlier model — and the groundwork the from-scratch Terse family builds on. A 2B ternary-weight LLM fine-tuned from Microsoft BitNet b1.58 on a single 4 GB laptop; on our own internal, self-run eval it lands third of twelve, ahead of every open-weight system we tested — including Qwen2.5-7B.

Explore Orchid 1.0
Internal eval · self-run Orchid 1.0
Claude 3.5 Sonnet
89.5
GPT-4o
89.2
Orchid 1.0 · 2B
87.9
BitNet b1.58 · 2B
84.2
Qwen2.5 · 7B
78.4

Our own internal, self-run eval (semantic-similarity scoring) — a relative comparison, not an official or standard NLP benchmark.


One project, three names

ternative — the project, and its in-house x86/AVX2 ternary inference engine.

terse — the from-scratch family of ternary-weight models (their own architecture).

orchid — Orchid 1.0, our earlier model: a fine-tune of Microsoft BitNet b1.58.


What we build

Two model lines, one
engine to run them.


Out in the open

Code, weights, paper —
all public.

micro-terse
Terse Micro model & training code — built from scratch in pure PyTorch
github.com ↗
Hugging Face · Orchid 1.0
Model card, GGUF weights & LoRA adapter
huggingface.co ↗
ternative engine
C++17 / CUDA inference engine — source, releases & build instructions
github.com ↗
Zenodo · Technical paper
DOI 10.5281/zenodo.20452163 — Orchid 1.0, archived & citable
zenodo.org ↗
FLOSS/fund
Support continued open development of Ternative, Terse & Orchid
floss.fund ↗