Classification & Scoring

Support chats: JEV scores, humans label, code decides

EN▶ videoClassification & ScoringHuman-reviewed entry
💬

Cover is generated by this site (not the creator's media).

WHAT IT DOES

Three jobs are kept separate: the chat model talks, JEV judges, and ordinary code decides what happens next.

Support conversations become sentiment probabilities, which are then compared against human labels so the review threshold can be chosen from data rather than taste.

Reported by the creator on X.

Summary written by this site

FROM THE ORIGINAL POST
Now I want to test Jev on support chats: chat → sentiment probabilities → compare with human labels → choose a review threshold. The chat model talks. Jev judges. Code decides what happens next. — Aman · Original post
JEV ROLE
Scoring Score #sentiment#benchmark#demo
WHERE IT SITS
HOW IT WORKS
INPUTa support conversation
JEVScoring
OUTPUTsentiment probabilities + a review decision
WHAT THE POST SHOWED
No observable metrics recorded. Metrics are only published when they can be tied to a source link.
SOURCE
↗ Original post
Source checked

Evidence is a status, not a ranking. It says how the claim can be checked — not whether the case is good.

RELATED CASES