← Back to Pulse PULSE. brief
I Tested Jev on 12 Real Use Cases. My Honest Thoughts.

I Tested Jev on 12 Real Use Cases. My Honest Thoughts.

Nate Herk | AI Automation16 min2026-09-19 ▶ Watch on YouTube
What this video is
⚡ a 16-minute video, readable in 60 seconds

This video reviews Jev, a new "System One Model" launched by TypeSafe AI and built by Diogo Almeida, who previously helped build the instruction-following methods behind ChatGPT at OpenAI, using a training method called Reinforcement Learning for Calibrated Decisions (RLCD). Unlike conversational LLMs, Jev skips string generation entirely and instead outputs fast yes/no, category, or score decisions, and the presenter tests it across 12 use cases including email triage, X-post labeling, and Bitcoin paper trading. Benchmarks show Jev roughly matching some models and trailing others on raw accuracy while running dramatically faster and cheaper at scale, though it has a small context window and can't write, explain, or handle math, dates, images, or video. The takeaway is that Jev fits high-volume classification and routing workflows, not general-purpose reasoning or chat.

[01:06] TypeSafe AI launched Jev, its first "System One Model," built by Diogo Almeida, who previously helped build ChatGPT's instruction-following methods at OpenAI, trained via Reinforcement Learning for Calibrated Decisions (RLCD).
Key takeaways
+ 9 more takeaways
  • [05:47] Classifying 1,000 emails across 7 rules took Jev about 70 seconds for 9 cents unparallelized; after switching to bigger parallel payloads, the same job finished in 6 seconds for the same 9 cents [06:30].
  • [06:11] The identical 1,000-email job on GPT-5.6 "Luna" took 5 minutes and cost 62 cents, about 46 times slower and 12 times pricier than Jev [08:19].
  • [07:25] Jev's invoice/receipt classification of the 1,000 emails finished in 4 seconds for 5 cents, flagging 237 "yes" and 763 "no."
  • [08:00] On the 5-level sponsor-fit scoring, 942 of the 1,000 emails scored the lowest tier (1) and none qualified as a "strong fit."
  • [00:27] In a Bitcoin paper-trading test, "JEV Trader" was down 2.82% of a $1,000 account after 1h24m, with 50 finished trades, 0 wins, and an average trade of -$0.56.
  • [14:57] Across 4,026 decisions, Jev's running cost was $0.091 total ($1.95/day), versus Sol at $159.73/day, Opus5 at $399.32/day, and Fable5.1 and Astra each at $798.63/day, per OpenRouter list prices.
  • [12:54] The recommended safeguard is to build a 100-example golden dataset and test Jev against Opus and Sol for accuracy versus cost before trusting it in production.
  • Current status is waiting in cash with $973.
  • Email dashboard showing results after sorting 1,000 emails in 4 seconds for $0.
Jump to a moment
Their links, sorted & clickable
🏛️ Communities & courses1My FREE resourcesskool.com
🛠️ Tools they use2FREE First Client SOPapp.aiautomationsociety.aiFREE MONTH voice to textget.glaido.com
📢 Sponsored / affiliate1Code NATEHERK for 10% off VPS (annual plan)hostinger.com
💼 Sponsorship & business1Sponsorship Inquiries
🌐 Find them3LinkedInlinkedin.comX / Twitterx.comInstagraminstagram.com
← Back to Pulse Dashboard
Was this brief useful?