Do we really needAn LLM for everything?
Is this a writing problem or a decision problem?
First, the basics
What's an LLM?
A Large Language Model is AI that has read a huge amount of text, so it can write like a person. You ask in plain words, and it writes an answer back.
- Writes, explains, summarises
- Answers in paragraphs
- Handles almost any question
- Takes a few seconds each time
A client says the export button does nothing and their report is due at 3pm. Is this a bug?
It sounds like a bug. The export button used to work, and now nothing happens when they click it, which points to something broken rather than a missing feature. Because they have a deadline at 3pm today, I'd treat it as high priority. It may be worth checking whether a recent release changed the export code…
Helpful, but a person still has to read it and decide what to do.
And the new one
What's Jev?
A different kind of AI model, made by TypeSafe. It doesn't write. It chooses.
If an LLM sits an essay exam, Jev sits a multiple-choice exam. And it tells you how confident it is in each answer.
“The export button does nothing. Our report is due at 3pm.”
What kind of request is this?
A client writes in
“Since this morning the export button does nothing when I click it. Our report is due at 3pm today.”
What should happen next?
One message. Four decisions.
- Bug or feature?bug · feature · question · unclear
- How urgent?urgent · high · normal · low
- Which part of the system?api · web · mobile · auth · infra
- Does a person need to see it?yes · no
None of those answers needs a paragraph.
A chat model
“Tell me what you think.”
- ✓Reading the message…
- ✓Reckoning…
- ✓Triaging…
- ✓Thinking…
- ✓Weighing it up…
- ✓Writing the answer…
Several steps and a few seconds later: a paragraph someone has to read, then act on.
A decision model · Jev
“Choose from these options.”
- ✓Choosing… ~0.5 s
One of our options, plus how sure it is. Code can act on it directly.
Real response · jev-1.13.0
What Jev sent back for the export button.
Every answer is one of the options we gave it. Nothing to parse, nothing to clean up. When api and web are this close, our rules label the area “unknown” rather than guess.
Type
Urgency
Area
Needs a person?
Give each part the job it's good at.
Before reaching for an LLM, ask which of these four the problem really is.
What we built
From client ticket to GitHub issue, without anyone triaging.
- Step 1ClientRaises a task in ClickUp
- Step 2WebhookClickUp tells our service
- Step 3 · ~0.5sJevType, urgency, area, needs a person. One call.
- Step 4Our rulesConfident enough to act?
- Step 5ActionTag + priority in ClickUp. Bugs become GitHub issues.
For the engineers
The AI part is about fifteen lines.
The rest is ordinary backend work: the webhook and its signature, retries, duplicate events, and mapping answers onto ClickUp and GitHub.
const { answers } = await jev.systemOne({ state: { title, description }, questions: { type: choice("What kind of request?", { bug, feature_request, question, unclear, }), urgency: choice("How urgent?", { urgent, high, normal, low, }), area: choice("Which part of the system?", AREAS), needs_human: noul("A person should read this first."), }, }); if (answers.type.confidence < 0.7) // → needs-review
The best part: it knows when it doesn't know.
The export button does nothing since this morning. Report due at 3pm.→ priority high · GitHub issue
It'd be great if we could export to Excel too.→ tagged · no GitHub issue
We were charged twice. We want a refund or we'll cancel.→ a person decides
Something about the app feels off lately.→ too vague · a person decides
Above the confidence line, the software acts. Below it, a person picks it up. We choose where the line sits: 70% today.
Live
- A complaint from a clientSomeone from HR or Media types it, not us
- A feature requestShould be tagged, with no GitHub issue
- Something vagueShould go to a person
$ npm start
Export button does nothing
Test run · 20 sample tickets · real Jev
Fast, cheap, and it sent the right ones to a person.
Tickets we wrote ourselves, not real client data. A small test, not a benchmark. 4 of the 20 went to a person: two vague ones, a refund dispute and a wrong invoice total.
It's not for everything.
Good fit
- Which team owns this?
- How urgent is it?
- Which category fits?
- Does this need a person?
- Which of these is most relevant?
Use something else
- Write the email back to the client
- Explain why the bug happens
- Design the architecture
- Write requirements or test cases
- Have a conversation
Not just engineering
The same idea at other desks.
“Forgot my laptop password” → IT
Route internal questions to the right place automatically.
Tag 50 posts in seconds. Only the doubtful ones get a second look.
Suggest how much testing a change needs. People still judge if it works.
Jev scores what needs attention today. Claude summarises only those five.
Two minutes · your turn
Every time X arrives, I look at it, pick from the same options, then do Y.
What's one decision you make at least five times a week? Shout it out.
Take this home
GitHub issue opened · 625 ms · $0.000034