60×
cheaper per answer
per 1,000 answers, when the smaller model passes your quality test.

We move each AI job to the right model, often a small one on your own hardware, and give it only the context it needs. The quality bar is agreed up front, and you keep the code.






Prior individual experience and employers of people we’ve trained. Not modalis clients, partners or endorsements.
60×
per 1,000 answers, when the smaller model passes your quality test.
2.5×
to write the same 500-word answer.
Yours
In your AWS, Google Cloud or Azure account, on rented H100s or A100s, or on the Macs you already own.
Cost per answer compares list prices for a frontier model (Claude Opus 5.5, GPT-5.6 Sol) with gpt-oss-20b on AWS Bedrock, for a retrieval answer of about 5,300 tokens in and 400 out; gpt-oss reasons before it answers, so real output can run longer. Speeds are for one request at a time: frontier APIs from Artificial Analysis medians; the Mac Studio M5 Max from public llama.cpp benchmarks with a 7B model at 4-bit; the H100 from an independent vLLM run of gpt-oss-20b. Figures as of September 2026.
Five steps, one senior team. The figures below are an example.
01 Audit
We map every model call your products and teams make: what it costs, how fast it is, how good it is, and what data it sends out.
You getA ranked savings plan with a measured baseline.
$35,000 a month, mapped call by call
02 Agree
Quality, cost, speed and data rules, written down and signed before we change anything.
You getA signed delivery charter.
Delivery charterSupport assistant
03 Build
Most calls go to a small model on your own hardware or cloud. The hard ones still reach a frontier model. Each call carries only the context it needs.
You getThe new setup, running in your cloud.
04 Measure
We measure the new setup against your current one, on real work, before anything changes for good.
You getA before-and-after report on real work.
05 Hand over
Models, code, tests and runbooks, and a team trained to run and extend them.
You getEverything, owned by your team.
Estimate your savings
How we’re paid
For select founding engagements, the implementation fee is due only after the written acceptance test passes. The test is agreed before any work starts.
We build our own product the way we build for clients: a small model on your Mac that keeps one page per project, so your next AI chat can start from it.
It depends on the job. We test it on your real examples against the quality bar we agree first. Work that needs a frontier model keeps one.
No. Models can run in your existing cloud account, on hardware you own, or both. We recommend what fits your data rules and your volume.
The call is free. We look at the AI you run or plan to run, what it costs, and where data rules or speed hold you back. Then we say what may be worth a closer look.
The discovery call is free. We scope any deeper work before it starts. The proposal lists each fee and outside cost, and how we measure the result.
You do. We deploy in your cloud and your repositories and hand over the code, models, tests and runbooks. You keep control after the engagement ends.
We work inside your cloud and use your access controls. Some work can run on hardware you own. We do not use your data to train models for anyone else.
Keep them where they are the best choice. We find the calls that don’t need them, move those to cheaper or private models, and cut the context each call carries.
Every change has a written test on real examples, agreed before we build. A change ships only when it meets that bar.
A free first conversation about the AI you run or plan to run: what it costs, how fast it is, and what data it sends out.