We move each AI job to the right model, often a small one on your own hardware, and give it only the context it needs. The quality bar is agreed up front, and you keep the code.

Track record

Delivered within
  • Nike
  • Mr. Cooper
  • Levi's
Trained professionals from
  • Salesforce
  • Meta
  • IBM
  • Microsoft
  • Samsung
  • HSBC
  • General Motors

Prior individual experience and employers of people we’ve trained. Not modalis clients, partners or endorsements.

Less spend. Faster answers.
Private when it matters.

60×

cheaper per answer

per 1,000 answers, when the smaller model passes your quality test.

2.5×

faster answers

Frontier API7.4 s
A Mac Studio in your office5.6 s
A rented H100 in your cloud2.9 s

to write the same 500-word answer.

Yours

models that run inside your walls

In your AWS, Google Cloud or Azure account, on rented H100s or A100s, or on the Macs you already own.

Sources

Cost per answer compares list prices for a frontier model (Claude Opus 5.5, GPT-5.6 Sol) with gpt-oss-20b on AWS Bedrock, for a retrieval answer of about 5,300 tokens in and 400 out; gpt-oss reasons before it answers, so real output can run longer. Speeds are for one request at a time: frontier APIs from Artificial Analysis medians; the Mac Studio M5 Max from public llama.cpp benchmarks with a 7B model at 4-bit; the H100 from an independent vLLM run of gpt-oss-20b. Figures as of September 2026.

How an engagement works.

Five steps, one senior team. The figures below are an example.

01 Audit

Find where the money goes.

We map every model call your products and teams make: what it costs, how fast it is, how good it is, and what data it sends out.

You getA ranked savings plan with a measured baseline.

$18.4kSupport assistant$9.2kDocuments$6.1kSearch$1.3kDrafts

$35,000 a month, mapped call by call

02 Agree

Agree the bar first.

Quality, cost, speed and data rules, written down and signed before we change anything.

You getA signed delivery charter.

Delivery charterSupport assistant

Quality
Today’s answers on 500 real tickets
Cost
Under one cent a ticket
Speed
Under two seconds
Data
Stays in your cloud
Signed before work starts
Agreed

03 Build

Send each job to the right model.

Most calls go to a small model on your own hardware or cloud. The hard ones still reach a frontier model. Each call carries only the context it needs.

You getThe new setup, running in your cloud.

Every request82% small model, in your cloud18% frontier model
Context per call3,000 tokens12,000

04 Measure

Hold the bar. Lower the bill.

We measure the new setup against your current one, on real work, before anything changes for good.

You getA before-and-after report on real work.

$41$9Per 1,000 tickets
3.8 s1.1 sAnswer time
91%92%Quality on the test set

05 Hand over

Then it’s yours.

Models, code, tests and runbooks, and a team trained to run and extend them.

You getEverything, owned by your team.

  1. 01Models and weightsin your account
  2. 02Code and serving setupin your repository
  3. 03Tests and runbooksyours to run
  4. 04A trained teamto run and extend it
Estimate your savings

Estimate your savings

What could your AI cost?

$13,560saved a month · $162,720 a year · 68% of today’s bill
check it with us A rough estimate from your numbers, not a quote. It assumes a smaller model costs about a tenth as much per call. The audit measures your real ones.

How we’re paid

Pay after the test passes.

For select founding engagements, the implementation fee is due only after the written acceptance test passes. The test is agreed before any work starts.

Acceptance test
Written and agreed before the build.
Implementation fee
Due after the test passes.
Performance fee
Separate, capped, and measured live where agreed.

The people you meet build the system.

Glyph portrait of Kukesh Kodess

Kukesh Kodess

co-founderlinkedin
Glyph portrait of Mike Fuller

Mike Fuller

co-founderlinkedin
our own product

Context your AI can pick up.

We build our own product the way we build for clients: a small model on your Mac that keeps one page per project, so your next AI chat can start from it.

whatprivate context layerrunslocally, on your Macreleasein development
Before you book

What buyers ask first.

01

Will a smaller model be good enough?

It depends on the job. We test it on your real examples against the quality bar we agree first. Work that needs a frontier model keeps one.

02

Do we need our own hardware?

No. Models can run in your existing cloud account, on hardware you own, or both. We recommend what fits your data rules and your volume.

03

What happens on the 25-minute call?

The call is free. We look at the AI you run or plan to run, what it costs, and where data rules or speed hold you back. Then we say what may be worth a closer look.

04

What does it cost?

The discovery call is free. We scope any deeper work before it starts. The proposal lists each fee and outside cost, and how we measure the result.

05

Who owns the code, models, and data?

You do. We deploy in your cloud and your repositories and hand over the code, models, tests and runbooks. You keep control after the engagement ends.

06

How do you handle our data?

We work inside your cloud and use your access controls. Some work can run on hardware you own. We do not use your data to train models for anyone else.

07

We already use ChatGPT or Claude. Why do we need you?

Keep them where they are the best choice. We find the calls that don’t need them, move those to cheaper or private models, and cut the context each call carries.

08

What if quality drops?

Every change has a written test on real examples, agreed before we build. A change ships only when it meets that bar.

Next step

Find out what your AI could cost.

Free discovery call
Twenty-five minutes on your AI spend.

A free first conversation about the AI you run or plan to run: what it costs, how fast it is, and what data it sends out.

what we discuss
  • 01your current AI use and monthly spend
  • 02where cost, speed, or data rules hold you back
  • 03what could move to a smaller or local model
25 minutes
Google Meet
America/Toronto