SYSTEM ONE MODEL INTELLIGENCE DIRECTORY|Community workflows and reports
Independent site

Independent community reports · Updated September 20, 2026

Jev with GPT, Claude, Gemini, Grok, Kimi and DeepSeek

See how people use Jev for game decisions, Claude context compaction, email screening and browser actions. Each report brings together original posts and documentation, including failures and unanswered questions.

6 model reports · 24 source notes

What Jev returns
API REFERENCE
Choice
Selects from options supplied by the application.
Score
Evaluates a question against a supplied rubric.
Noul
Returns a probability for a statement.

Workflows and feedback by model

A runtime handoff, a coding assistant and a comparison test are different ways two models can appear together. The reports explain which relationship the sources actually describe.

OpenAI3 sources

Jev + GPT / Codex

GPT revising Jev game prompts, plus a separate Codex desktop experiment. What the users reported, including failed runs.

Read report
Anthropic3 sources

Jev + Claude

A plugin scores old tool calls before pruning context. Its implementation and a user report show both the approach and its limits.

Read report
Google10 sources

Jev + Gemini

Two published comparisons cover business emails and app reviews. Gemini leads on the reported quality checks; Jev costs less in those runs.

Read report
xAI4 sources

Jev + Grok

A 1,000-post selection workflow and a separate action-router setup, with source links and no inferred speed or cost claims.

Read report
Moonshot AI3 sources

Jev + Kimi

A measured email handoff sends uncertain cases to Kimi K3. A separate developer uses Kimi to build a Jev-powered comment filter.

Read report
DeepSeek4 sources

Jev + DeepSeek

Developers report intent-recognition changes and a train-search demo. The browser project documents how decisions and text generation are separated.

Read report

Measurements in context

The cited authors measured these workloads. We preserve their sample sizes and distinguish a first-pass duration from the full pipeline. The site has not independently rerun the experiments.

Business-email classification

Nikhil Mudholkar: 1,565 emails across 10 categories, including 1,201 real emails and 364 AI-written examples.

Results and limitations

Labeling 1,000 app reviews

Rahul Kumar’s Column Race: sentiment, topic, bug and churn labels for 1,000 Android reviews. Both lanes use batches of 20 at concurrency 8.

Results and limitations

Email screening with a Kimi K3 fallback

Hassan: 100 emails, evenly split between legitimate and fraudulent examples. Jev sends 31 cases below 95% confidence to Kimi K3.

Results and limitations

From the original discussions

All 24 sources

These are editorial summaries of the linked posts, not direct quotations.

paulwei (@coolish)

Published:

Source checked:

X.com

[4] Astra revises Jev prompts after a failed game run

The first Ascension 10 attempt reached floor 6. The author says Astra then iterated on Jev prompts and the pair reached floor 17, while describing Jev as weaker at the game than Astra.

A follow-up from the same experiment, not independent corroboration. Reaching floor 17 is not a reported completed run.

Jev / GPTRead original

Alex Volkov (@altryne)

Published:

Source checked:

X.com

[7] A user report of Claude context reduction

Volkov reports reducing a Claude session from nearly one million tokens to about 86,000 in roughly one second using fast-jev-compaction.

Individual experience, not a controlled benchmark or a measure of downstream task accuracy. The timestamp is the edited-post time shown by X.

Jev / ClaudeRead original

Dan McAteer (@daniel_mac8)

Published:

Source checked:

X.com

[17] Ranking X posts with Jev and Grok Bot

McAteer describes retrieving 1,000 agent-related posts with the X API, using Jev for pairwise selection, and having Grok Bot surface five tips.

A first-hand workflow description. No model version, comparison count, latency or cost breakdown is supplied.

Jev / GrokRead original

Hassan (@nutlope)

Published:

Source checked:

X.com

[19] Jev first-pass email screening with Kimi K3 fallback

Hassan tests 50 legitimate and 50 fraudulent emails. Jev classifies them in 1.42 seconds; 31 cases below a 95% confidence threshold go to Kimi K3. The full pipeline takes 16 seconds, gets 96 of 100 correct and costs about $0.07.

Author-reported demonstration, not independently rerun here. The stated cost split is $0.003 for Jev and $0.068 for Kimi. The 1.42-second figure excludes the Kimi stage.

Jev / KimiRead original

For a closer look at the division of work, read the Jev workflow patterns: confidence-based handoffs, context pruning and browser action selection.