Tiny Judgments, Priced Three Ways: Jev, Laya, and OpenAI's Decisions API
· AgenticLabs India
· AgenticLabs India
Every automated workflow burns full LLM calls on tiny judgments — classify this, route that, approve or block. Three new products sell the judgment without the prose: TypeSafe's Jev, open-weight Laya, and OpenAI's Decisions API. Here is what each costs, how fast each answers, and which one fits a small shop.
Picture the most common automation in a small business: a customer message arrives, and something has to decide what it is. Sales enquiry, support complaint, or spam? Then route it to the right person. That is one judgment — one label, one confidence number. It takes a human about five seconds.
But the way most businesses automate it today, they rent the heaviest machinery available: a full large language model call, which writes a paragraph of reasoning and explanation just so the workflow can parse out the single word "sales". The paragraph gets thrown away. You paid for the prose; you kept the label.
That is the hidden tax. Thousands of tiny judgments a month, each one billed like a conversation.
In the last month, three products have attacked exactly this problem, all from the same insight: most automation decisions are "System 1" thinking — fast, pattern-matching, one-shot — not "System 2" deliberation. They sell you the judgment without the paragraph:
(One housekeeping note before we compare: the decision-model Laya lives at convaiinnovations/laya on Hugging Face. It is a different project from aayushch/laya, a notification command center that happens to share the name. Make sure you land on the right one.)
Jev is the newcomer that defined the category. You send it a "state" — text or JSON describing the situation — plus questions defined in code, and it returns only structured values: a probability that a condition holds, one pick from a fixed option list (up to 255 options), or a position on an ordinal scale. No prose, no explanation, no code. It evaluates every question in a call independently in a single parallel pass — the model is non-autoregressive, so it never writes tokens one by one — and it was trained with a method called RLCD (reinforcement learning for calibrated decisions) that optimizes for honest probabilities, not fluent writing.
The company behind it is TypeSafe AI, founded by Diogo Almeida, an ex-OpenAI researcher, which came out of stealth in mid-September 2026 with $40M in seed funding.
Two facts make Jev's economics concrete. First, the published price: $0.042 per 1M input tokens, with output tokens free and unmetered (TypeSafe's docs, corroborated by its listing on the Vercel AI Gateway at about $0.04/1M). Second, the speed: TypeSafe claims 70–500 ms end-to-end (a vendor claim — treat vendor claims as vendor claims), while independent measurements by SCAND clocked roughly 700 ms for a 1,000-token request against 2,300–4,400 ms for GPT-6-class calls on the same workload, and community testers report 115–217 ms on small calls. Either way: several times faster than a full LLM call, at a small fraction of the price.
Jev already has the liveliest ecosystem of the three, which is the strongest signal about a three-week-old product. Real things people have built: jev-ultrafast, an open-source browser agent built with Browser Use where Jev picks the next action from a numbered menu each step (github.com/browser-use/jev-ultrafast); jevmail, an open-source Gmail triage tool that its author reports sorted 1,000 emails in about a minute for roughly $0.03 (github.com/fazlerocks/jevmail); HA-Jev, a Home Assistant integration exposing typed answers as sensors; and CartShield, a small-business checkout fraud disposition tool. You can also reach Jev through Vercel's AI Gateway as typesafe-ai/jev.
Laya is the same idea with the opposite business model. It is a 421M-parameter decision model — a ModernBERT-large encoder plus a small decision head — released under the Apache 2.0 licence by independent researcher Nandakishor M of Convai Innovations, three days after Jev launched. It answers the same typed questions (pick from a list, score on a scale, probability a condition holds) in a single forward pass, never generating text. Three checkpoints are published on Hugging Face: an English one (512-token context), a multilingual one covering 100+ languages with a 1,024-token context, and one fine-tuned specifically on typed-decision workflows.
The pitch is simple: the weights are free, so you pay only your own compute — and your data never leaves your building. There is no waitlist and no vendor gate; availability is a function of your own hardware. pip install laya gets you the model, and laya-serve starts a local server with a Jev-compatible API shape, so the same question definitions work without the cloud. On hardware you already own, the marginal cost of each decision is effectively zero.
On speed, the project's own figures put a single decision at about 33 ms on a Tesla T4; independent testers report 18–40 ms on consumer GPUs, around 13 ms on an Apple M3 Max via the MLX port, and 500+ ms on CPU-only machines. Batching is nearly free on a GPU. In plain terms: on any machine with a graphics card, Laya answers in the time it takes to blink.
The community has already put it to work locally: Ghosthand (github.com/ameerhmz/ghosthand), a push-to-talk voice assistant and computer-use agent for macOS that runs fully on Apple Silicon with no cloud APIs; laya-ultrafast, a port of the browser agent that swaps the Jev API call for a local Laya pass; Clara, a private on-PC assistant using Laya as its router; and Acurast's announcement of running Laya on decentralized smartphone compute.
OpenAI's entry is not a new model — it is a constrained interface over the existing one. The Decisions API (POST /v1/decisions) gives GPT-6 Luna a pre-defined set of options and forbids it from writing anything else. Three question types: predicate (the probability a condition holds), choice (one option from a fixed set, with per-option probabilities and a confidence value), and score (a rating on ordered levels, returned as a probability-weighted average). It accepts text and images, though images must arrive as inline base64 data — hosted links and uploaded file IDs are refused. There is a Playground at platform.openai.com/decisions for trying it by hand.
It is also the youngest: announced at DevDay on September 29 as a limited preview, in public beta since October 6, with general availability promised "in the coming weeks." The beta pricing is published: $0.10 per 1M input tokens, with no charge for output tokens or cache reads/writes (OpenAI's docs, corroborated by Cellcog's docs-sourced write-up and AIWeekly's beta coverage). On speed, OpenAI claims roughly 10× faster than the standard Responses API — but as multiple write-ups note, that figure ships without a published benchmark, and as of this writing no independent measurements exist. Treat it as a vendor claim until someone measures it.
On trust and compliance, OpenAI offers Zero Data Retention and HIPAA eligibility for eligible customers, but processing happens in the US and Europe only — no India/APAC region, which matters if your customers' data should stay close to home.
A fair note on scope: the API is three days into public beta, so there are no production deployments, case studies, or community builds to point at yet — only OpenAI's own docs examples (a damaged-product photo check, a billing-complaint router, a bug-severity scorer) and third-party explainer write-ups. Compare that with Jev and Laya, which each grew a working open-source ecosystem in about three weeks.
One ecosystem footnote: there is also Kev, an Apache-2.0 open-weight decision model from Jared Palmer (released September 20) designed as a Jev API drop-in — worth knowing as proof the category is bigger than these three, but out of scope for this comparison.
Let us make it concrete. Suppose your shop routes 10,000 customer messages a month, and each message plus the question costs about 500 input tokens — 5M tokens a month. None of the three charge for output in decision mode, so the arithmetic is simple:
At this volume, all three cost less than a cup of chai a month. The money only matters when you scale: at a million decisions a month, Jev costs about $21 and the Decisions API about $50 (Cellcog published the same figures), while Laya stays flat at your compute cost. That is the real pricing story — hosted options charge per judgment; the open model charges per machine.
Speed, honestly labelled:
Accuracy and calibration, from the evidence available: an independent study published October 5 tested Jev against Laya on 11 decision benchmarks (7,283 cases) and found Jev winning 9 of 11, by 10.8 to 46 percentage points. Laya's own authors are candid about why: the base English checkpoint is weak at zero-shot and expects to be fine-tuned on your workflow — the fine-tuned typed-decisions checkpoint is a different animal. Laya also ships over-confident out of the box and needs per-workflow temperature calibration, and the authors advise keeping option lists under about 20 choices. Jev, for its part, calibrates honestly across groups but not per individual answer, and like all three it is weak at arithmetic, counting, and date comparison — it reads instructions literally and can be confidently wrong within a valid type. The Decisions API has no published head-to-head accuracy numbers at all, against anything. The only accuracy test that ultimately matters is on your own messages, which is why the use case below ends with a scoring step.
Privacy and lock-in:
Maturity and capacity risk, stated plainly: Jev is in early access behind a waitlist, and independent testers report real overload errors under load. Laya is community-maintained open source — no SLA, no support line. The Decisions API is three days into public beta with GA "coming weeks" and its beta pricing subject to change. None of these is the boring, five-year-old SaaS you already trust. Budget accordingly: start with one low-stakes workflow, not your order pipeline.
Here is the smallest valuable thing you can hand to a decision model: sorting the WhatsApp messages that pile up while you work. Every Indian small business knows the pile — a mix of genuine sales enquiries, support follow-ups, and spam, all demanding attention in the order they arrived instead of the order they deserve. A decision model can read each message and return one label — sales, support, or spam — with a confidence number. You look at the sales ones first, the spam never.
Below is a follow-along you can genuinely do this weekend, no developer needed for the first five steps. You will test the same messages on all three products, score them yourself, and finish knowing which one deserves your business.
Copy ten recent customer messages from your WhatsApp Business chat into a notes app — or use these five stand-ins to start, then swap in your own:
1. "Hi, do you deliver to Whitefield? Need 50 gift boxes by Friday."
2. "My order #4521 hasn't arrived, it's been 6 days. Please check."
3. "CONGRATULATIONS!! You have WON a lottery. Click here to claim your prize."
4. "What is the price for the 2kg variant? And is it eggless?"
5. "Bhaiya, kal jo parcel bheja tha woh damage aaya. Replacement milega?"Message 5 is there on purpose: a Hindi/Hinglish support message. Multilingual handling is one of the real differences between these products, and you want to see it before you commit.
Expected result: ten messages in a note, each one a real message your business actually received.
Jev is in early access, so start at the TypeSafe console and request access; developers can also reach it through Vercel's AI Gateway. The question you define is the whole product — and it is genuinely copy-paste. In the console's playground, create one question with exactly these fields:
Question name: intent
Question type: choice
Instructions: Decide what kind of customer message this is. Pick exactly one of the options below.
Options: sales, support, spamThen paste each of your ten messages in as the "state" (the situation the question is asked about) and run it. Jev answers with the chosen label and a confidence number — the intent question comes back answered as sales with something like 0.97 confidence. No paragraph, no explanation.
For a developer wiring this into code later, the documented request shape is a POST to /v1/systemone carrying a model name (for example jev-latest), your message as state, and each question defined by its instructions and criteria — one entry per question key. (TypeSafe's own docs describe the endpoint, and the API reference documents the full schema.)
Expected result: ten labels with ten confidence numbers, each returned in well under a second.
Open the Decisions Playground (public beta — you will need an OpenAI API account). Create a choice question with the same three options — sales, support, spam — and paste the same ten messages. The API returns the chosen option with per-option probabilities and a confidence value.
Expected result: ten labels with confidence numbers, directly comparable to Step 2's. Note the response time yourself — nobody independent has published measurements yet, so your stopwatch is currently state of the art.
This step is the one the other two cannot offer. On Laya's Hugging Face page (which hosts a live demo), or locally with pip install laya and laya-serve (a Jev-compatible local server, so the same question definitions work without the cloud), run the same ten messages. Two things to try deliberately: use the multilingual checkpoint for message 5 (the English checkpoint is documented to stay confident while degrading off English — route language first), and keep your option list short (the authors advise under ~20 options).
Expected result: ten labels, answered locally — and a direct feel for the "no cloud involved" difference. If Laya's labels look shaky on your messages, that is the documented fine-tuning gap, not a bug: the base checkpoint expects specialization.
Make a simple sheet — paper works — with these columns:
message | jev label | jev confidence | decisions label | decisions confidence | laya label | my labelFill in your own label last, then count agreements per product. Pay special attention to two things: messages where a product was confident and wrong (the most expensive kind of mistake), and how each product handled the Hindi message. Ten messages will not give you statistical significance, but they will tell you which product reads your customers correctly — and the independent benchmarks suggest Jev will lead on many-option accuracy while Laya needs tuning to compete.
Expected result: a scored sheet and a winner for your message style — evidence, not marketing.
Pick the product that scored best on your messages and give it one boring, bounded job: every evening, run the day's unanswered WhatsApp messages through the winning question, and sort them into three lists — sales, support, spam. Act only on labels with confidence at or above 0.9; hand-check everything below it. That threshold rule is the entire quality system. This is the same "start read-only, review everything" posture as our five mundane workflows — and if you want the routine wired into your inbox or CRM instead of run by hand, that is exactly the kind of workflow automation we build.
Expected result: tomorrow morning you open three short lists instead of one long pile — and you know, from your own scoring, how much to trust each label.
Three questions decide it:
1. May customer messages leave your premises? If the answer is no — medical, financial, or simply your own comfort — Laya is the only option of the three. It is also the only one with no per-message bill and no vendor who can change the price or the terms. The price you pay is setup effort and fine-tuning.
2. Do you need it working this week with no setup? Then you want hosted, which means Jev or the Decisions API. Jev is cheaper per token ($0.042 vs $0.10 per 1M input), has independent speed measurements, and a livelier builder ecosystem — but sits behind a waitlist and has shown real overloads. The Decisions API has OpenAI's brand and compliance apparatus (Zero Data Retention and HIPAA for eligible customers) plus image inputs, but is three days into beta with unverified speed and no production track record.
3. How many judgments a month? Under ~50,000 decisions a month, the hosted options cost a few dollars a month at published prices, and Laya's setup effort is not worth it. Past a few hundred thousand a month, Laya's flat compute cost pulls ahead decisively — the per-judgment meter is exactly what open weights eliminate.
Two special cases: if your messages are multilingual, Laya's multilingual checkpoint (100+ languages) is the most explicit answer — though Jev and Luna handle multiple languages too, and you should test message 5 from the use case on all three. If you need to judge images (a damaged-product photo, a blurry invoice scan), the Decisions API is currently the only one of the three that takes image input at all.
A fair comparison says what is still raw:
Start with the weekend test in the use case above. Ten messages, three products, one sheet. The right answer for your shop is the one that reads your customers best — measured, not marketed.
It is an AI model that answers only structured judgments — a probability, a pick from a fixed list, or a score — instead of generating text. Because it never writes prose, it answers in milliseconds and costs a fraction of a full language-model call. Think of it as the difference between asking someone "write an essay about this message" and "which of these three boxes does it go in?"
At roughly 500 input tokens per message: about $0.21 a month on Jev ($0.042 per 1M input tokens, output free), about $0.50 a month on OpenAI's Decisions API ($0.10 per 1M input tokens in beta, no output charges), and effectively $0 on Laya plus your own compute. Those are published early-access/beta prices and can change — the method matters more than the numbers.
Test it before you trust it — that is why the use case includes a Hindi message. Laya's multilingual checkpoint explicitly covers 100+ languages; its English checkpoint is documented to stay confident while degrading off English, so route language first. Jev and OpenAI's Luna handle multiple languages, but multilingual accuracy is exactly the kind of thing you verify on your own messages in Step 5.
For the weekend test: no. Jev's console, OpenAI's Decisions Playground, and Laya's Hugging Face demo Space all let you paste messages and read labels by hand. Turning the winner into an automatic daily routine — pulling messages, calling the API, sorting the lists — is developer work, or the kind of workflow automation a consultant wires up once and hands over.
It depends on which you pick. With Laya, messages never leave your machine — the strongest privacy story of the three, with the caveat that you own the security setup. With Jev, data travels to TypeSafe's US servers (a DPA is available; customer data is not used for training per their docs). With OpenAI's Decisions API, data travels to OpenAI's US/Europe infrastructure, with Zero Data Retention available only for eligible customers. Match the product to your comfort and your compliance needs.
On published evidence: an independent October 2026 study of 11 decision benchmarks found Jev beating Laya on 9 of 11 tasks, while Laya's fine-tuned checkpoint narrows the gap substantially. The Decisions API has no published head-to-head accuracy numbers yet. But "most accurate on benchmarks" and "most accurate on your customers' Hinglish" are different questions — run the ten-message test in this post and let your own sheet decide.
Book a free consultation. We'll show you which of your workflows AI can take over first.