AI automationdecision modelscost comparison

Tiny Judgments, Priced Three Ways: Jev, Laya, and OpenAI's Decisions API

· AgenticLabs India

Every automated workflow burns full LLM calls on tiny judgments — classify this, route that, approve or block. Three new products sell the judgment without the prose: TypeSafe's Jev, open-weight Laya, and OpenAI's Decisions API. Here is what each costs, how fast each answers, and which one fits a small shop.

The hidden tax on every tiny judgment

Picture the most common automation in a small business: a customer message arrives, and something has to decide what it is. Sales enquiry, support complaint, or spam? Then route it to the right person. That is one judgment — one label, one confidence number. It takes a human about five seconds.

But the way most businesses automate it today, they rent the heaviest machinery available: a full large language model call, which writes a paragraph of reasoning and explanation just so the workflow can parse out the single word "sales". The paragraph gets thrown away. You paid for the prose; you kept the label.

That is the hidden tax. Thousands of tiny judgments a month, each one billed like a conversation.

In the last month, three products have attacked exactly this problem, all from the same insight: most automation decisions are "System 1" thinking — fast, pattern-matching, one-shot — not "System 2" deliberation. They sell you the judgment without the paragraph:

  • Jev, a hosted decision API from TypeSafe AI (launched September 15, 2026, currently in early access with a waitlist), that returns probabilities and picks from fixed option lists instead of text.
  • Laya, an open-weight decision model from Convai Innovations (released September 18, 2026, Apache 2.0 licence), that runs entirely on your own machine — no cloud, no API key, no per-call bill.
  • OpenAI's Decisions API, a dedicated endpoint announced at DevDay 2026 on September 29 and released in public beta on October 6, that constrains the GPT-6 Luna model to developer-defined questions with pre-defined answers.

(One housekeeping note before we compare: the decision-model Laya lives at convaiinnovations/laya on Hugging Face. It is a different project from aayushch/laya, a notification command center that happens to share the name. Make sure you land on the right one.)

Three answers to the same question

Jev (TypeSafe AI): the hosted decision API

Jev is the newcomer that defined the category. You send it a "state" — text or JSON describing the situation — plus questions defined in code, and it returns only structured values: a probability that a condition holds, one pick from a fixed option list (up to 255 options), or a position on an ordinal scale. No prose, no explanation, no code. It evaluates every question in a call independently in a single parallel pass — the model is non-autoregressive, so it never writes tokens one by one — and it was trained with a method called RLCD (reinforcement learning for calibrated decisions) that optimizes for honest probabilities, not fluent writing.

The company behind it is TypeSafe AI, founded by Diogo Almeida, an ex-OpenAI researcher, which came out of stealth in mid-September 2026 with $40M in seed funding.

Two facts make Jev's economics concrete. First, the published price: $0.042 per 1M input tokens, with output tokens free and unmetered (TypeSafe's docs, corroborated by its listing on the Vercel AI Gateway at about $0.04/1M). Second, the speed: TypeSafe claims 70–500 ms end-to-end (a vendor claim — treat vendor claims as vendor claims), while independent measurements by SCAND clocked roughly 700 ms for a 1,000-token request against 2,300–4,400 ms for GPT-6-class calls on the same workload, and community testers report 115–217 ms on small calls. Either way: several times faster than a full LLM call, at a small fraction of the price.

Jev already has the liveliest ecosystem of the three, which is the strongest signal about a three-week-old product. Real things people have built: jev-ultrafast, an open-source browser agent built with Browser Use where Jev picks the next action from a numbered menu each step (github.com/browser-use/jev-ultrafast); jevmail, an open-source Gmail triage tool that its author reports sorted 1,000 emails in about a minute for roughly $0.03 (github.com/fazlerocks/jevmail); HA-Jev, a Home Assistant integration exposing typed answers as sensors; and CartShield, a small-business checkout fraud disposition tool. You can also reach Jev through Vercel's AI Gateway as typesafe-ai/jev.

Laya (Convai Innovations): the open-weight decision model

Laya is the same idea with the opposite business model. It is a 421M-parameter decision model — a ModernBERT-large encoder plus a small decision head — released under the Apache 2.0 licence by independent researcher Nandakishor M of Convai Innovations, three days after Jev launched. It answers the same typed questions (pick from a list, score on a scale, probability a condition holds) in a single forward pass, never generating text. Three checkpoints are published on Hugging Face: an English one (512-token context), a multilingual one covering 100+ languages with a 1,024-token context, and one fine-tuned specifically on typed-decision workflows.

The pitch is simple: the weights are free, so you pay only your own compute — and your data never leaves your building. There is no waitlist and no vendor gate; availability is a function of your own hardware. pip install laya gets you the model, and laya-serve starts a local server with a Jev-compatible API shape, so the same question definitions work without the cloud. On hardware you already own, the marginal cost of each decision is effectively zero.

On speed, the project's own figures put a single decision at about 33 ms on a Tesla T4; independent testers report 18–40 ms on consumer GPUs, around 13 ms on an Apple M3 Max via the MLX port, and 500+ ms on CPU-only machines. Batching is nearly free on a GPU. In plain terms: on any machine with a graphics card, Laya answers in the time it takes to blink.

The community has already put it to work locally: Ghosthand (github.com/ameerhmz/ghosthand), a push-to-talk voice assistant and computer-use agent for macOS that runs fully on Apple Silicon with no cloud APIs; laya-ultrafast, a port of the browser agent that swaps the Jev API call for a local Laya pass; Clara, a private on-PC assistant using Laya as its router; and Acurast's announcement of running Laya on decentralized smartphone compute.

OpenAI's Decisions API: the incumbent's answer

OpenAI's entry is not a new model — it is a constrained interface over the existing one. The Decisions API (POST /v1/decisions) gives GPT-6 Luna a pre-defined set of options and forbids it from writing anything else. Three question types: predicate (the probability a condition holds), choice (one option from a fixed set, with per-option probabilities and a confidence value), and score (a rating on ordered levels, returned as a probability-weighted average). It accepts text and images, though images must arrive as inline base64 data — hosted links and uploaded file IDs are refused. There is a Playground at platform.openai.com/decisions for trying it by hand.

It is also the youngest: announced at DevDay on September 29 as a limited preview, in public beta since October 6, with general availability promised "in the coming weeks." The beta pricing is published: $0.10 per 1M input tokens, with no charge for output tokens or cache reads/writes (OpenAI's docs, corroborated by Cellcog's docs-sourced write-up and AIWeekly's beta coverage). On speed, OpenAI claims roughly 10× faster than the standard Responses API — but as multiple write-ups note, that figure ships without a published benchmark, and as of this writing no independent measurements exist. Treat it as a vendor claim until someone measures it.

On trust and compliance, OpenAI offers Zero Data Retention and HIPAA eligibility for eligible customers, but processing happens in the US and Europe only — no India/APAC region, which matters if your customers' data should stay close to home.

A fair note on scope: the API is three days into public beta, so there are no production deployments, case studies, or community builds to point at yet — only OpenAI's own docs examples (a damaged-product photo check, a billing-complaint router, a bug-severity scorer) and third-party explainer write-ups. Compare that with Jev and Laya, which each grew a working open-source ecosystem in about three weeks.

One ecosystem footnote: there is also Kev, an Apache-2.0 open-weight decision model from Jared Palmer (released September 20) designed as a Jev API drop-in — worth knowing as proof the category is bigger than these three, but out of scope for this comparison.

The three-way tradeoff: cost, speed, and privacy

Let us make it concrete. Suppose your shop routes 10,000 customer messages a month, and each message plus the question costs about 500 input tokens — 5M tokens a month. None of the three charge for output in decision mode, so the arithmetic is simple:

  • Jev: 5M tokens × $0.042/1M = about $0.21 a month.
  • OpenAI Decisions API: 5M × $0.10/1M = about $0.50 a month.
  • Laya: $0 for the weights plus your own compute — effectively zero marginal cost on hardware you already own.

At this volume, all three cost less than a cup of chai a month. The money only matters when you scale: at a million decisions a month, Jev costs about $21 and the Decisions API about $50 (Cellcog published the same figures), while Laya stays flat at your compute cost. That is the real pricing story — hosted options charge per judgment; the open model charges per machine.

Speed, honestly labelled:

  • Jev: vendor claim 70–500 ms; independent measurements 115–700 ms depending on workload — several times faster than a full LLM call.
  • Laya: about 33 ms on a T4 GPU (project figure), 18–40 ms measured on consumer GPUs, ~13 ms on Apple Silicon — but 500+ ms on CPU-only machines. Your hardware decides.
  • Decisions API: "about 10× faster than the Responses API" and "a few hundred milliseconds" — both OpenAI's words, both unverified by anyone independent yet.

Accuracy and calibration, from the evidence available: an independent study published October 5 tested Jev against Laya on 11 decision benchmarks (7,283 cases) and found Jev winning 9 of 11, by 10.8 to 46 percentage points. Laya's own authors are candid about why: the base English checkpoint is weak at zero-shot and expects to be fine-tuned on your workflow — the fine-tuned typed-decisions checkpoint is a different animal. Laya also ships over-confident out of the box and needs per-workflow temperature calibration, and the authors advise keeping option lists under about 20 choices. Jev, for its part, calibrates honestly across groups but not per individual answer, and like all three it is weak at arithmetic, counting, and date comparison — it reads instructions literally and can be confidently wrong within a valid type. The Decisions API has no published head-to-head accuracy numbers at all, against anything. The only accuracy test that ultimately matters is on your own messages, which is why the use case below ends with a scoring step.

Privacy and lock-in:

  • Jev: your data travels to TypeSafe's servers (US West Coast only, which adds latency from India and makes EU data-transfer rules your problem; TypeSafe offers a DPA and says it does not train on customer data, with zero-retention for enterprise plans only). Closed weights, hosted only — you rent, you do not own. No SOC 2 is mentioned in the public docs.
  • Laya: your data never leaves your machine. Apache 2.0, self-hosted, with a Jev-compatible local server — the least lock-in of any AI product we have covered. The catch: you own the security hardening and the fine-tuning.
  • Decisions API: your data travels to OpenAI (US + Europe processing only). Zero Data Retention and HIPAA eligibility exist but are eligibility-gated, not default.

Maturity and capacity risk, stated plainly: Jev is in early access behind a waitlist, and independent testers report real overload errors under load. Laya is community-maintained open source — no SLA, no support line. The Decisions API is three days into public beta with GA "coming weeks" and its beta pricing subject to change. None of these is the boring, five-year-old SaaS you already trust. Budget accordingly: start with one low-stakes workflow, not your order pipeline.

The practical use case: triage your WhatsApp enquiries for pennies

Here is the smallest valuable thing you can hand to a decision model: sorting the WhatsApp messages that pile up while you work. Every Indian small business knows the pile — a mix of genuine sales enquiries, support follow-ups, and spam, all demanding attention in the order they arrived instead of the order they deserve. A decision model can read each message and return one label — sales, support, or spam — with a confidence number. You look at the sales ones first, the spam never.

Below is a follow-along you can genuinely do this weekend, no developer needed for the first five steps. You will test the same messages on all three products, score them yourself, and finish knowing which one deserves your business.

Step 1: Collect ten real messages

Copy ten recent customer messages from your WhatsApp Business chat into a notes app — or use these five stand-ins to start, then swap in your own:

1. "Hi, do you deliver to Whitefield? Need 50 gift boxes by Friday."
2. "My order #4521 hasn't arrived, it's been 6 days. Please check."
3. "CONGRATULATIONS!! You have WON a lottery. Click here to claim your prize."
4. "What is the price for the 2kg variant? And is it eggless?"
5. "Bhaiya, kal jo parcel bheja tha woh damage aaya. Replacement milega?"

Message 5 is there on purpose: a Hindi/Hinglish support message. Multilingual handling is one of the real differences between these products, and you want to see it before you commit.

Expected result: ten messages in a note, each one a real message your business actually received.

Step 2: Ask Jev the same question ten times

Jev is in early access, so start at the TypeSafe console and request access; developers can also reach it through Vercel's AI Gateway. The question you define is the whole product — and it is genuinely copy-paste. In the console's playground, create one question with exactly these fields:

Question name: intent
Question type: choice
Instructions: Decide what kind of customer message this is. Pick exactly one of the options below.
Options: sales, support, spam

Then paste each of your ten messages in as the "state" (the situation the question is asked about) and run it. Jev answers with the chosen label and a confidence number — the intent question comes back answered as sales with something like 0.97 confidence. No paragraph, no explanation.

For a developer wiring this into code later, the documented request shape is a POST to /v1/systemone carrying a model name (for example jev-latest), your message as state, and each question defined by its instructions and criteria — one entry per question key. (TypeSafe's own docs describe the endpoint, and the API reference documents the full schema.)

Expected result: ten labels with ten confidence numbers, each returned in well under a second.

Step 3: Ask OpenAI's Decisions API the same question

Open the Decisions Playground (public beta — you will need an OpenAI API account). Create a choice question with the same three options — sales, support, spam — and paste the same ten messages. The API returns the chosen option with per-option probabilities and a confidence value.

Expected result: ten labels with confidence numbers, directly comparable to Step 2's. Note the response time yourself — nobody independent has published measurements yet, so your stopwatch is currently state of the art.

Step 4: Ask Laya — on your own machine

This step is the one the other two cannot offer. On Laya's Hugging Face page (which hosts a live demo), or locally with pip install laya and laya-serve (a Jev-compatible local server, so the same question definitions work without the cloud), run the same ten messages. Two things to try deliberately: use the multilingual checkpoint for message 5 (the English checkpoint is documented to stay confident while degrading off English — route language first), and keep your option list short (the authors advise under ~20 options).

Expected result: ten labels, answered locally — and a direct feel for the "no cloud involved" difference. If Laya's labels look shaky on your messages, that is the documented fine-tuning gap, not a bug: the base checkpoint expects specialization.

Step 5: Score them like an exam

Make a simple sheet — paper works — with these columns:

message | jev label | jev confidence | decisions label | decisions confidence | laya label | my label

Fill in your own label last, then count agreements per product. Pay special attention to two things: messages where a product was confident and wrong (the most expensive kind of mistake), and how each product handled the Hindi message. Ten messages will not give you statistical significance, but they will tell you which product reads your customers correctly — and the independent benchmarks suggest Jev will lead on many-option accuracy while Laya needs tuning to compete.

Expected result: a scored sheet and a winner for your message style — evidence, not marketing.

Step 6: Turn the winner into a daily routine

Pick the product that scored best on your messages and give it one boring, bounded job: every evening, run the day's unanswered WhatsApp messages through the winning question, and sort them into three lists — sales, support, spam. Act only on labels with confidence at or above 0.9; hand-check everything below it. That threshold rule is the entire quality system. This is the same "start read-only, review everything" posture as our five mundane workflows — and if you want the routine wired into your inbox or CRM instead of run by hand, that is exactly the kind of workflow automation we build.

Expected result: tomorrow morning you open three short lists instead of one long pile — and you know, from your own scoring, how much to trust each label.

Which one fits your shop?

Three questions decide it:

1. May customer messages leave your premises? If the answer is no — medical, financial, or simply your own comfort — Laya is the only option of the three. It is also the only one with no per-message bill and no vendor who can change the price or the terms. The price you pay is setup effort and fine-tuning.

2. Do you need it working this week with no setup? Then you want hosted, which means Jev or the Decisions API. Jev is cheaper per token ($0.042 vs $0.10 per 1M input), has independent speed measurements, and a livelier builder ecosystem — but sits behind a waitlist and has shown real overloads. The Decisions API has OpenAI's brand and compliance apparatus (Zero Data Retention and HIPAA for eligible customers) plus image inputs, but is three days into beta with unverified speed and no production track record.

3. How many judgments a month? Under ~50,000 decisions a month, the hosted options cost a few dollars a month at published prices, and Laya's setup effort is not worth it. Past a few hundred thousand a month, Laya's flat compute cost pulls ahead decisively — the per-judgment meter is exactly what open weights eliminate.

Two special cases: if your messages are multilingual, Laya's multilingual checkpoint (100+ languages) is the most explicit answer — though Jev and Luna handle multiple languages too, and you should test message 5 from the use case on all three. If you need to judge images (a damaged-product photo, a blurry invoice scan), the Decisions API is currently the only one of the three that takes image input at all.

The honest limits, in plain words

A fair comparison says what is still raw:

  • Everything here is weeks old. Jev launched September 15, Laya September 18, the Decisions API entered public beta October 6. You are an early adopter no matter which you pick; none has the five-year track record of the SaaS tools you already run.
  • Jev is capacity-constrained. Early access, a waitlist, documented overload errors under real load, and US-West-Coast-only serving. Fine for experiments; think twice before it sits in your order pipeline.
  • Laya demands homework. The base checkpoint is weak until fine-tuned on your workflow, ships over-confident until you calibrate it, and wants short option lists. If "fine-tune a 421M-parameter model" sounds like a weekend project you will never start, the hosted options are more honest for you.
  • The Decisions API is a beta. Pricing, docs, and behaviour can change before general availability; the 10× speed claim has no independent verification; processing stays in the US and Europe.
  • None of them explain themselves. A decision model returns a label and a number, not a reason. That is the point — and it means your quality system is the confidence threshold plus human review, not the model's eloquence. All three can be confidently wrong; keep the 0.9 rule from the use case.
  • Prices quoted are published early-access/beta prices (Jev via TypeSafe's docs, Decisions API via OpenAI's beta announcement) and can change. Re-check before you build a budget on them — the per-10,000-decision arithmetic above is a comparison method, not a quote.

Start with the weekend test in the use case above. Ten messages, three products, one sheet. The right answer for your shop is the one that reads your customers best — measured, not marketed.

Questions, answered

It is an AI model that answers only structured judgments — a probability, a pick from a fixed list, or a score — instead of generating text. Because it never writes prose, it answers in milliseconds and costs a fraction of a full language-model call. Think of it as the difference between asking someone "write an essay about this message" and "which of these three boxes does it go in?"

Ready to leave the mundane work to AI?

Book a free consultation. We'll show you which of your workflows AI can take over first.