Built automatically with zero data, it beat Jev in Doom and Snake
Send us the JSON you send Jev.In minutes, get your own 1 MB model.
For games and rule-based decisions, it beats Jev from day one. For text, it beats Jev clearly once it learns from your decisions.
5-day free trial, no card. Then $20 a month, flat.
- Doom, zero data
- 45 kills to 39
- 2 deaths to Jev's 5 at the same decision rate. Won 7 of 13 against its default style, 12 of 13 against its aggressive one.
- Snake, zero data
- 42.0 vs 40.0
- food at equal steps. Never died; ahead in 12 of 20 games, level in 2.
- Bank messages
- 91.7% vs 86.6%
- after about 400 reviewed cases; about 95% with your history. Close to Jev on day one.
- Size
- < 1 MB
- per decision model. Text models share one 34 MB reader.
Answers
A bank support model you can try: type a message of your own, or flip the account fields and watch the rules change the decision.
How it works
Start with no data. Use it. Make it better when you like.
Three steps to your own model, then about a minute per improvement. No past cases needed to start.
- 01
Paste your Jev request
The JSON you already send Jev: one real example of your state and your questions, as they are. Add your rules in plain words if you like.
POST /v1/build{ "state": { "message": "I want my money back…","amount": 412 },"questions": { "action": { "type": "choice", … } },"rules": "Never refund over £250." } - 02
Your model is ready in minutes
It's built, trained and checked against your instructions and rules on situations it never saw. Nothing to do in between. Review how it decides any time (optional): correct any example and it retrains.
- Built from your JSONdone
- Traineddone
- Checked: 98% match your rulesready
- Review 20 examplesoptional
- 03
Use it
Change the URL and the model name, and keep sending the same JSON. Every answer says where it came from and which rules applied.
- api.typesafe.ai · "jev-latest"+ api.canonopylabs.com · "refunds@latest""action": "escalate-senior","blocked_by_rules": ["never-refund…"],"route": "act"
Then
Let it learn from your decisions
Answer the cases your model was unsure about, or send your past cases, and retrain in about a minute. You see a before and after, and every retrain is your call. On real bank messages: about 86% on day one, 91.7% after about 400 reviewed cases, about 95% with the bank's history.
Rules found in your history
Escalate refunds over 300
True in 97% of the 610 past cases it applies to. Only the rules you confirm are used.
ConfirmReject
Improve your model
Answer a few of these
These 5 decisions in your unsure queue are the ones your model was least sure about. Each answer teaches your next retrain the most.
Answer them, then retrain
Results
Beats Jev in games from day one. On text, once it learns from your decisions.
The game players below were built automatically from only what Jev is given: the state, the questions and the rules. No past cases, no recorded play.
Doom
13 games against recorded live-Jev games, both sides deciding at the same rate
45 kills to 39
2 deaths to 5 · won 7 of 13 · zero data
Against Jev's default style. Against its aggressive style: 45 kills to 21, 2 deaths to 11, won 12 of 13. Built in about 4 minutes.
Snake
Jev's own demo game, 20 games stopped at the same step
42.0 vs 40.0
food per game · never died · zero data
Ahead of Jev's default strategy in 12 of 20 games, level in 2, behind in 6. Jev crashed once.
Bank support
3,080 real customer messages, nine next steps, five hard rules
91.7% vs 86.6%
after about 400 reviewed cases
Close to Jev on day one with no data: 85.5% across 10 automatic builds. About 95% with the bank's history.
On text, your decisions make the difference.
The same 3,080 bank messages, each routed to one of nine next steps under five hard rules. Same messages, same rules, both sides.
Hover a bar for how it was measured.
Hover a bar for how it was measured.
Hover a bar for how it was measured.
Canonopy's $0.78 is computer time if you run it yourself. Hosted by us, it's the flat $20 a month.
Hover a bar for how it was measured.
View as a table
| Measure | Canonopy, with history | After ~400 reviewed cases | Day one, no data | Jev, few-shot | Jev, zero-shot |
|---|---|---|---|---|---|
| Accuracy | ~95% | 91.7% | 85.5% | 86.6% | 83.9% |
| Calibration error | 0.006 | – | – | 0.044 | 0.063 |
| Time per decision (p50) | 38.9 ms | – | – | 177 ms | 216 ms |
| Cost per million decisions | $0.78 | – | – | $158.61 | $55.58 |
Canonopy's $0.78 is computer time if you run it yourself. Hosted by us, it's the flat $20 a month.
Two more real datasets, with history
Consumer complaints
Real complaints to the US CFPB, each routed at intake
77.7%Canonopy, against Jev 65.0% few-shot, 62.9% zero-shot
The categories are what consumers picked on the complaint form, so part of the lead is matching that form.
Moderation
Civil Comments: allow, review or remove
62.7%Canonopy, against Jev 57.9% zero-shot, 57.0% few-shot
Ours catches more harmful comments; Jev leaves more civil ones alone. A simple model on the same text reader ties ours here.
The honest footnote. Doom is one automatic build: 13 games per style on one map, scored against recorded games of live Jev with both sides deciding every 500 ms. Snake is the best of four automatic builds on crashes (three of the four ate more than Jev): 20 games against Jev's recorded games on the same seeds. Against Jev's greedy strategy, Jev ate more (40.8 to 35.6) but crashed in 13 games, ours in none. On text, a model built with no data is close to Jev on short messages (84.6–86.8% against 86.6%), not ahead, and well below it on long complaint narratives; it pulls clearly ahead once it learns from your decisions. The reviewed cases were answered with their true answers, as a person reviewing would, starting from a no-data model made from hand-written examples (88.6% before any cases); answers that come mostly from a backup model will likely help less. Rules on what a message is about are enforced whenever there's a real chance they apply, and unsure cases go to review.
Messages are the banking77 test split (PolyAI, CC BY 4.0). It's public, so either model may have seen it before. Accuracy is against a written routing policy: some of Jev's misses were defensible calls, and counting those as right the lead with history is still about 5–7 points. The account details attached to each message were assigned at random for the test, so real account data is still untested. Our times and costs were measured on a laptop CPU; Jev's over HTTPS at list price, network included. Complaints and moderation are public datasets too, so either model may have seen them.
What you need to bring
Nothing but your Jev request. Your decisions make it better.
The same request works whatever your decisions read: messages, data and state, or a game.
Messages
tickets, emails, chats, claims, reviews
You bring
Your Jev request
Close to Jev on day one for short messages; give it past cases or answer a few reviews to pull clearly ahead. The answers you already record are enough. It reads about 100 languages; accuracy measured in 51.
Data and state
transactions, account fields, readings, schedules
You bring
Your Jev request
Numbers go into the model as numbers, so it can tell £4,999 from £5,001. Rules on your fields are enforced on every decision, whatever the model thinks.
Games and simulations
opponents, NPCs, agents in a simulator
You bring
Your game's Jev request
Send the state and the moves, get the next move in milliseconds, on our servers or inside your game loop. Beat Jev in Doom and Snake from day one; recorded play, when you have it, makes it better.
Mixed, like a bank case with account details and a message? The same request covers both.
Where this came from. We were building minds for intelligent NPCs. When Jev launched we tested them against it, found we had built System One models, and are offering them first as this API.
Use it with your LLM
Let the LLM talk. Let your model decide.
Language models are great at writing and reading between the lines. Put the rule-bound decision somewhere that can't be talked out of your rules.
Support bot
Your model picks the action; your LLM writes the reply around it. The bot can't promise a refund your rules don't allow.
See the pattern
Guardrails
Before an LLM's proposed action runs, check it. blocked_by_rules tells you which rule stopped it, in words you wrote.
See the pattern
Agent decision layer
Give your agent one tool: decide. It gets calibrated answers and a route (act, review, ask) instead of guessing.
See the pattern
Pricing
Try it free. Then $20 a month, flat.
One price per workspace, however many decision models you make. A 5-day free trial that starts when you first use it (your first model or first decision). No card. No tokens, no per-call bill. Run it on our servers or yours.
Flat. Unlimited decision models, decisions and retraining, up to 30 builds a month. Everything included.
- Unlimited decision models
- Unlimited decisions
- Unlimited retraining and reports
- Up to 30 builds a month
- Rules found in your history
- Outcomes and the unsure queue
- Download to run it yourself
- Your models are yours to keep
Fair use: the hosted endpoint has a per-second rate limit (50 requests per second by default). A workspace can start 5 builds a day and 30 a month, 3 during the free trial (a build is a new model or a rule change), and 20 retrains a day. Run it yourself for any volume.
What a million decisions costs elsewhere
Per-token pricing grows with every call and every word in the prompt.
- Hosted Jev, zero-shot
- $55.58
- Hosted Jev, few-shot
- $158.61
- Canonopy, hosted
- $20
1,323 input tokens per request, at list price
3,776 input tokens per request, at list price
a month, flat, for all your decision models, whether you make a thousand decisions or a hundred million
Jev figures are from our banking77 benchmark run, billed per input token.
Get started
Your own model, in minutes.
Create a workspace with your email and paste the JSON you send Jev. Your model is ready in minutes.
A 5-day free trial from your first model or first decision, no card. Then $20 a month, flat, with unlimited decision models, decisions and retraining and up to 30 builds a month. Your models are yours to keep.
Try it free