SevenX tech talk · Backend teamJev × ClickUp × GitHub

Do we really needAn LLM for everything?

Is this a writing problem or a decision problem?

From: clientTicket 0001
Route to: ?
Write a reply Make a decision

First, the basics

What's an LLM?

A Large Language Model is AI that has read a huge amount of text, so it can write like a person. You ask in plain words, and it writes an answer back.

You know theseChatGPTClaudeGemini
  • Writes, explains, summarises
  • Answers in paragraphs
  • Handles almost any question
  • Takes a few seconds each time
A chat with an LLMessay answer

A client says the export button does nothing and their report is due at 3pm. Is this a bug?

It sounds like a bug. The export button used to work, and now nothing happens when they click it, which points to something broken rather than a missing feature. Because they have a deadline at 3pm today, I'd treat it as high priority. It may be worth checking whether a recent release changed the export code…

Helpful, but a person still has to read it and decide what to do.

And the new one

What's Jev?

A different kind of AI model, made by TypeSafe. It doesn't write. It chooses.

You give itThe information, a question, and the possible answers
It gives backOne of those answers, and how sure it is

If an LLM sits an essay exam, Jev sits a multiple-choice exam. And it tells you how confident it is in each answer.

“The export button does nothing. Our report is due at 3pm.”

What kind of request is this?

Bug100%
Feature request0%
Question0%
Unclear0%
~0.5 secondsa fraction of a centno paragraph to read

A client writes in

ClickUp · new task09:14

“Since this morning the export button does nothing when I click it. Our report is due at 3pm today.”

What should happen next?

One message. Four decisions.

  1. Bug or feature?bug · feature · question · unclear
  2. How urgent?urgent · high · normal · low
  3. Which part of the system?api · web · mobile · auth · infra
  4. Does a person need to see it?yes · no

None of those answers needs a paragraph.

A chat model

“Tell me what you think.”

  1. ✓Reading the message…
  2. ✓Reckoning…
  3. ✓Triaging…
  4. ✓Thinking…
  5. ✓Weighing it up…
  6. ✓Writing the answer…

Several steps and a few seconds later: a paragraph someone has to read, then act on.

A decision model · Jev

“Choose from these options.”

  1. ✓Choosing… ~0.5 s
Bug100% confident

One of our options, plus how sure it is. Code can act on it directly.

Real response · jev-1.13.0

What Jev sent back for the export button.

625 ms 807 tokens in $0.000034

Every answer is one of the options we gave it. Nothing to parse, nothing to clean up. When api and web are this close, our rules label the area “unknown” rather than guess.

Ticket 00014 questions · 1 call

Type

bug100%
feature0%

Urgency

high99%
urgent1%

Area

api53%
web47%

Needs a person?

yes31%

Give each part the job it's good at.

ActsCodeThings we know exactly. It's a bug, so set the priority.
DecidesJevSmall judgments, fast. Bug or feature? How urgent?
WritesLLMReasoning and writing. Explain the bug, draft the reply.
ApprovesPeopleImportant or unclear calls. Refund disputes, vague tickets.

Before reaching for an LLM, ask which of these four the problem really is.

What we built

From client ticket to GitHub issue, without anyone triaging.

  1. Step 1ClientRaises a task in ClickUp
  2. Step 2WebhookClickUp tells our service
  3. Step 3 · ~0.5sJevType, urgency, area, needs a person. One call.
  4. Step 4Our rulesConfident enough to act?
  5. Step 5ActionTag + priority in ClickUp. Bugs become GitHub issues.

For the engineers

The AI part is about fifteen lines.

The rest is ordinary backend work: the webhook and its signature, retries, duplicate events, and mapping answers onto ClickUp and GitHub.

const { answers } = await jev.systemOne({
  state: { title, description },
  questions: {
    type: choice("What kind of request?", {
      bug, feature_request, question, unclear,
    }),
    urgency: choice("How urgent?", {
      urgent, high, normal, low,
    }),
    area: choice("Which part of the system?", AREAS),
    needs_human: noul("A person should read this first."),
  },
});

if (answers.type.confidence < 0.7) // → needs-review

The best part: it knows when it doesn't know.

The export button does nothing since this morning. Report due at 3pm.→ priority high · GitHub issue
Bug
It'd be great if we could export to Excel too.→ tagged · no GitHub issue
Feature
We were charged twice. We want a refund or we'll cancel.→ a person decides
Review
Something about the app feels off lately.→ too vague · a person decides
Review

Above the confidence line, the software acts. Below it, a person picks it up. We choose where the line sits: 70% today.

Live

  1. A complaint from a clientSomeone from HR or Media types it, not us
  2. A feature requestShould be tagged, with no GitHub issue
  3. Something vagueShould go to a person

$ npm start

ClickUp · client requestsjust now

Export button does nothing

BugHigh
🤖 Jev triage: bugbug (100% confident) · urgency high (99%)
GitHub issue opened · 625 ms · $0.000034

Test run · 20 sample tickets · real Jev

Fast, cheap, and it sent the right ones to a person.

20/20Bug, feature, question or unclear: all correct
8/10Urgency right. Both misses were one level off.
523msMedian per ticket. Slowest: 699 ms.
<0.1¢For all 20 tickets together ($0.000675)

Tickets we wrote ourselves, not real client data. A small test, not a benchmark. 4 of the 20 went to a person: two vague ones, a refund dispute and a wrong invoice total.

It's not for everything.

Good fit

  • Which team owns this?
  • How urgent is it?
  • Which category fits?
  • Does this need a person?
  • Which of these is most relevant?

Use something else

  • Write the email back to the client
  • Explain why the bug happens
  • Design the architecture
  • Write requirements or test cases
  • Have a conversation

Not just engineering

The same idea at other desks.

DeskHRHR“When does my probation end?” → employment
“Forgot my laptop password” → IT

Route internal questions to the right place automatically.

DeskMDMediacampaign? audience? → needs brand review?

Tag 50 posts in seconds. Only the doubtful ones get a second look.

DeskQAQAchange → smoke · targeted · full

Suggest how much testing a change needs. People still judge if it works.

DeskMGManagers40 threads, tickets, PRs → top 5

Jev scores what needs attention today. Claude summarises only those five.

Two minutes · your turn

Every time X arrives, I look at it, pick from the same options, then do Y.

What's one decision you make at least five times a week? Shout it out.

Which tests should I run? Who should review this PR? Where does this request go? Does this need attention today?

Take this home

Before calling an LLM, ask: is this writing or deciding?

LLM writes · Jev decides · Code acts · People approve

Questions?

Jev · typesafe.aiBackend team · SevenX

Keys

→ Space
Next
← Backspace
Previous
O
Overview
N
Speaker notes
F
Fullscreen
? Esc
Help

Speaker notes