
Shape a multilingual inventory before training.
A measured release benchmark showing the effect of retained Unicode bigrams on one FineWeb2 Chinese workload.
View repositoryOpen source lab / models + agents
We build reusable primitives, developer tools, and hosted services for the hard problems shared across AI products.
Showcase
Representative outputs from the current public projects. Each one solves a different shared problem beneath AI products.
Open projects
Each project tackles a difficult shared layer. Use one on its own, or combine them without adopting an entire stack.
Fast and faithful byte-pair encoding for training exact tokenizers over large, multilingual corpora.
A composable Python workspace for language-model research, training, reinforcement learning, and inference.
One local, OpenAI-compatible endpoint for supported coding plans and model providers.
Fast, local reports that show where your coding-agent tokens and estimated costs went.
About the lab
Most teams focus on the product they need to ship today. We focus on the foundations many products need, but few teams have the time to build well.
We build the tools behind AI products.
Open at the core. Local when preferred. Hosted when useful.01_reusableWe invest in clear boundaries and durable interfaces so the same foundation can support many different ideas.
02_openThe core stays inspectable and replaceable. You should be able to see the tradeoffs and choose how each piece runs.
03_adaptableFocused components adapt better than heavy abstractions. We keep the pieces small enough to evolve and strong enough to depend on.
Build with us
Explore the projects, use what fits, and help improve the layers beneath the next generation of AI products.