AIProduct engineering

Be Like Percy

Where Jev makes EosHQ faster: categorising requests, choosing Agents and recognising automation, with measured time and model-cost savings.

A retro microwave and chocolate bar connected by glowing paths to three decision tiles, one highlighted

In 1945, Percy L. Spencer, standing in front of a Raytheon radar device, noticed that the candy bar in his pocket had melted. Instead of being pissed off with the mess (I would be), he wondered how he could repurpose the effect for other uses. After testing his idea with popping popcorn and blowing up an egg in a colleague's face, he figured he could build something that would cook food without ever lighting a stove. Beginning with the dangerous-sounding name of "Radarange", that new device is today's microwave. We wouldn't use it for ALL our cooking, but it is the exact device we need for some of the things we want done.

Jev is something like that: a specialised tool that does certain jobs very well. We wouldn’t ask it to write an explanation or work through a complicated business problem. But give it a clear decision and a fixed set of options, and it can be remarkably quick. Not for all uses, but in the right use-case, Jev beats the pants off of LLMs. And TypeSafe has been a lot more creative than Raytheon in naming this new tech. 😊

So what's Jev? And how do we use it at EosHQ? First, about Jev: simply put, Jev is a decision engine. Consider the last time your team decided where to go for lunch. Lots of conversation, long sentences, preferences, heartburn (of the literal kind) and so on till you finally decided on Taco Bell. That's like talking to ChatGPT or Claude or Gemini - lots of words, with the decision hidden somewhere in there in ambiguous language. Jev isn't like that at all. It takes your problem as input, with a list of options as results, and returns one of those options - even just a yes or no. So basically, you get to Taco Bell before they run out of the Crispy Chicken Nuggets. General-purpose LLMs can make that choice too - Jev got us there faster.

We've been looking at that difference here at EosHQ, wanting to make Eos Chat conversations as fast as possible. Some messages need EosHQ to interpret records, work through a problem or write an explanation. Others simply need to choose from a few known options. That second group is where we now use Jev. Where we've used it, it's worked extremely well.

Start with the decision the business needs

Jev lets us give it a question with relevant information and a defined set of possible answers. For example, a ticket might belong to billing, technical support or sales. We give it the ticket body and the 3 options and bam, it returns a choice and confidence information. We still decide which choices are allowed, what each one means and what the application may do afterwards. But that choice went from 3 seconds per request to 455 ms. Quite the improvement, I'd say.

Let me describe the various situations we're using Jev in.

Putting a request in the right category The first use is classification. Like I said above, we use Jev when the request has a set of labels and asks us to choose one. One billing example used “billing, technical, or sales” and the text “I was charged twice for my subscription this month.” Jev correctly selected billing. From a business perspective, this is the decision you make before handing work to the appropriate team or process. So our implementation replaces the label-selection step. Jev chooses the category. Assignment and resolution remain separate steps, governed by their own rules.

In our live example, the recorded classification call fell from 3.07 seconds with another LLM to 0.455 seconds with Jev, an 85% reduction. Those measurements include some application overhead around the model request. We separately measured Jev’s network request at 0.251 seconds. This was one observed request per version. The earlier Gemini path also lost the supplied labels and generated a rationale, so the comparison includes a change in prompt and output behaviour.

Finding the right Agent The next use is choosing a published Agent. Imagine that your org has an Agent for reviewing overdue invoices and another for preparing a project risk digest. So a request to “review overdue invoices and prioritise collection follow-up” should reach the "overdue invoices" Agent. And “Show my overdue tasks” should not accidentally launch the "overdue invoices" Agent just because both requests contain the word “overdue.” In our tests, the Jev-enabled flow selected the correct Agent, including the option of "none" in some cases. Some decisions needed Gemini fallback. In EosHQ, the list contains only Agents that the user is allowed to access, so this is the classic case of a request leading to only one of multiple options - right up Jev's alley.

Finding the right automation model Another use is to recognize an automation request. “Show my overdue tasks” asks for a response immediately. But “Every Monday, prepare a summary of overdue tasks” asks for a recurring process. And “Whenever an invoice is created, check for a purchase order number” describes an event-triggered process. Jev handles this initial classification, with Gemini taking over when confidence is too low. That classification is only one step towards defining and approving the resulting automation. In our tests, the invoice-trigger example needed Gemini in all three repetitions and was slightly slower with Jev in front of it.

Time those decisions

For Agent selection and automation intent, we ran 72 synthetic tests through production decision functions in an isolated harness, using a synthetic catalogue of eligible Agents. The comparator was Gemini 3 Flash Preview. Each test (business function) had six scenarios, repeated three times with each provider. We included straightforward requests, partial matches, ambiguity and quoted instructions. The tests measured the model decision stage, excluding database access and the rest of the Chat response. The Agent and automation figures below measure that decision stage, including Gemini fallbacks. The classification observation includes application overhead; the skill-selection row is an earlier experiment. None measures a complete Chat response.

Function Specific replacement Gemini time Jev path, including fallback Time saved Gemini → Jev cost / 1,000 Cost saved / 1,000
Categorise a Request Select one supplied label, such as billing / technical / sales, for the provided text. 3,070 ms 455 ms 2,615 ms · 85.2% $1.0905 → $0.0150 $1.0755 · 98.6%
Select the right EosHQ Skill (earlier experiment) Choose the appropriate skill for supported task, PTO, calendar, time-entry, sentiment and classification requests. 2,636 ms 477 ms 2,159 ms · 81.9% $5.2863 → $0.3378 $4.9485 · 93.6%
Select the right Agent Given a user's message, select the right primary agent to respond with 3,380 ms 1,616 ms 1,764 ms · 52.2% $1.293 → $0.526 $0.767 · 59.3%
Detect the right Automation Type Recognise a scheduled, event-triggered or other automation request, or no automation 3,432 ms 925 ms 2,507 ms · 73.1% $1.366 → $0.234 $1.132 · 82.8%

When Jev’s confidence falls below our threshold, we ask Gemini to make the decision independently. Jev requests alone averaged 365 ms for Agent selection and 350 ms for automation detection, approximately 9–10× faster than the corresponding Gemini requests. But low-confidence results still required Gemini:

  • Agent selection: 6 of 18 Jev-enabled decisions fell back.
  • Automation detection: 3 of 18 fell back.

So the benefit depends on the kind of request, even within a function where the overall average looks good.

The money is real, but keep the scale visible

We also calculated model costs from the returned token counts and the providers’ published prices. At the rates used for this comparison, Gemini 3 Flash Preview costs $0.50 per million text input tokens and $3 per million output tokens, including thinking. Jev costs $0.042 per million input tokens, with output free. Here is what 1,000 repetitions of our measured decisions would cost at those rates. These are estimated model charges, excluding infrastructure and other Chat steps:

Decision Gemini estimate Jev estimate, including fallback Estimated saving
Classify text into supplied labels $1.091 $0.015 98.6%
Choose an existing Agent $1.293 $0.526 59.3%
Recognise an automation request $1.366 $0.234 82.8%

An 83% saving on one decision does not mean an 83% reduction in an application’s AI bill. And anyway, at modest volumes, the absolute amounts are small. But our interest was in reducing the time you spend waiting, with lower model costs as a welcome additional benefit.

One useful surprise: sometimes we need no model

Interesting side-note: Our re-engineering effort exposed a redundant step in our own implementation. We were using Jev to select a skill even when the application had already recognised a complete, supported request and knew its arguments. For those exact requests, we now select the skill directly, without either model making that selection. The earlier skill-selection experiment in the table therefore describes a step those requests now skip. That helped the complete classification Chat response move from 3.29 to 2.63 seconds after it was already using Jev. It also explains why a quarter-second model call can still sit inside a conversation that takes several seconds: permissions, context, data access, accounting and the final response all take time. There is more to improve, and making the fastest component faster will not necessarily change the part you are waiting for.

We will keep using Jev where a clearly defined decision benefits from it, while retaining the LLM for broader reasoning and generated responses. For your own processes, the useful starting point is to identify the small decisions people or software repeatedly make before the larger work begins. Bring one of those processes to an EosHQ conversation, and we can look at where the waiting actually happens. And bring a candy bar to see if you have the time to eat it while Jev (or the LLM) is working on your request. 😊

Pricing references: TypeSafe’s Jev model reference and Gemini pricing. Measurements were collected on September 28, 2026. These small samples describe our tests; they are not a production latency guarantee.

Back to the blog