Independent community reports · Updated September 20, 2026
Jev with GPT, Claude, Gemini, Grok, Kimi and DeepSeek
See how people use Jev for game decisions, Claude context compaction, email screening and browser actions. Each report brings together original posts and documentation, including failures and unanswered questions.
6 model reports · 24 source notes
- Choice
- Selects from options supplied by the application.
- Score
- Evaluates a question against a supplied rubric.
- Noul
- Returns a probability for a statement.
Sources:[1] TypeSafe AI
Workflows and feedback by model
A runtime handoff, a coding assistant and a comparison test are different ways two models can appear together. The reports explain which relationship the sources actually describe.
Jev + GPT / Codex
GPT revising Jev game prompts, plus a separate Codex desktop experiment. What the users reported, including failed runs.
Jev + Claude
A plugin scores old tool calls before pruning context. Its implementation and a user report show both the approach and its limits.
Jev + Gemini
Two published comparisons cover business emails and app reviews. Gemini leads on the reported quality checks; Jev costs less in those runs.
Jev + Grok
A 1,000-post selection workflow and a separate action-router setup, with source links and no inferred speed or cost claims.
Jev + Kimi
A measured email handoff sends uncertain cases to Kimi K3. A separate developer uses Kimi to build a Jev-powered comment filter.
Jev + DeepSeek
Developers report intent-recognition changes and a train-search demo. The browser project documents how decisions and text generation are separated.
Measurements in context
The cited authors measured these workloads. We preserve their sample sizes and distinguish a first-pass duration from the full pipeline. The site has not independently rerun the experiments.
Business-email classification
Nikhil Mudholkar: 1,565 emails across 10 categories, including 1,201 real emails and 364 AI-written examples.
Results and limitationsLabeling 1,000 app reviews
Rahul Kumar’s Column Race: sentiment, topic, bug and churn labels for 1,000 Android reviews. Both lanes use batches of 20 at concurrency 8.
Results and limitationsEmail screening with a Kimi K3 fallback
Hassan: 100 emails, evenly split between legitimate and fraudulent examples. Jev sends 31 cases below 95% confidence to Kimi K3.
Results and limitationsFrom the original discussions
All 24 sourcesThese are editorial summaries of the linked posts, not direct quotations.
paulwei (@coolish)
Published:
Source checked:
[4] Astra revises Jev prompts after a failed game run
The first Ascension 10 attempt reached floor 6. The author says Astra then iterated on Jev prompts and the pair reached floor 17, while describing Jev as weaker at the game than Astra.
A follow-up from the same experiment, not independent corroboration. Reaching floor 17 is not a reported completed run.
Alex Volkov (@altryne)
Published:
Source checked:
[7] A user report of Claude context reduction
Volkov reports reducing a Claude session from nearly one million tokens to about 86,000 in roughly one second using fast-jev-compaction.
Individual experience, not a controlled benchmark or a measure of downstream task accuracy. The timestamp is the edited-post time shown by X.
Dan McAteer (@daniel_mac8)
Published:
Source checked:
[17] Ranking X posts with Jev and Grok Bot
McAteer describes retrieving 1,000 agent-related posts with the X API, using Jev for pairwise selection, and having Grok Bot surface five tips.
A first-hand workflow description. No model version, comparison count, latency or cost breakdown is supplied.
Hassan (@nutlope)
Published:
Source checked:
[19] Jev first-pass email screening with Kimi K3 fallback
Hassan tests 50 legitimate and 50 fraudulent emails. Jev classifies them in 1.42 seconds; 31 cases below a 95% confidence threshold go to Kimi K3. The full pipeline takes 16 seconds, gets 96 of 100 correct and costs about $0.07.
Author-reported demonstration, not independently rerun here. The stated cost split is $0.003 for Jev and $0.068 for Kimi. The 1.42-second figure excludes the Kimi stage.
For a closer look at the division of work, read the Jev workflow patterns: confidence-based handoffs, context pruning and browser action selection.