JobBench: Aligning Agent Work with Human Will

Measuring agents by GDP alone asks how much of a human's job can be taken away.

JobBench asks how much of that job can be given back โ€” built on the work that experts across real-world professions actually want delegated to AI.

agent_01
Current leader
Claude Fable 5
Anthropic
Weighted score57.4%
0
Professions
0
Tasks
0
Criteria

In collaboration with

University of Washington
UC Santa Barbara
Stanford University
Carnegie Mellon University
University of Notre Dame
IBM Research
BakeAI
Michigan State University
UC Berkeley
Northwestern University
University of Chicago

Adopted by

JobBench has been adopted by Meta's Muse Spark 1.1 and Moonshot's Kimi K3

Muse Spark 1.1Meta
Kimi K3Moonshot
ยง 01 โ€” Why Human Will

Economics alone is not enough.

The conversation about AI in the workplace has been framed almost entirely in economic terms: what fraction of working hours can agents absorb? how much of GDP is exposed to automation? Benchmarks like OpenAI's GDPval inherit this framing by design โ€” they select tasks that represent economic value, and score agents on whether they can deliver the professional knowledge output.

We believe this framing, on its own, is not enough.

If agents are going to share the professional workplace with humans, the question is not only what work is most economically valuable to automate, but what work do the humans in that role actually want automated? This is a humanist problem. It treats the professional not as labor to be displaced, but as a collaborator whose judgment about their own craft matters โ€” and it is the premise JobBench is built on.

The economic question

GDPval

OpenAI

โ€œWhat fraction of a human's job is economically valuable to automate?โ€

The humanist question

JobBench

Ours

โ€œWhat work do the humans in that role actually want automated?โ€

Read the full essay
ยง 02 โ€” Rankings

Model leaderboard

Family
1
Claude Fable 5
57.4
2
Muse Spark 1.1
54.7
3
Kimi K3
54.3
04
Claude Opus 4.8
48.4
05
GPT-5.6 SOL
45.4
06
Claude Opus 4.7
44.5
07
GLM 5.2
43.4
08
GPT-5.5
38.3
09
Claude Sonnet 4.6
36.6
10
GPT-5.4
32.2
11
Gemini 3.5 Flash
31.5
12
GPT-5.2
26.6
13
Claude Sonnet 4.5
20.7
14
Gemini 3.1 Pro
15.9

All models run on the same harness โ€” OpenCode v1.14.18 โ€” with corresponding max reasoning effort. Grok 4.3 is used as the rubric judge. (Earlier runs used Grok 4.1 Fast, since retired by xAI, so scores may differ slightly from previously reported results.) The Grok judges are chosen mainly for cost: one full eval pass costs ~$2 with Grok 4.1 Fast and ~$20 with Grok 4.3. Claude Fable 5 runs with fallback to Claude Opus 4.8 on refusals.

ยง 03 โ€” Headroom

Far from saturation

GPT-5.4 โ€” Codex CLI
GDPvalsaturating
83.0
JobBench61 pts headroom
38.9
GPT-5.2 Codex
70.9/24.8
GPT-5.3 Codex
70.9/33.7
GPT-5.4 โ€” Codex CLI
83.0/38.9
Workload
JobBench over GDPval
Wall-clock per task2.40ร—
Tool calls per task1.40ร—
Trajectory lines1.40ร—
ยง 04 โ€” Methodology

From knowledge delivery to professional reasoning

ยง 05 โ€” Inside a task

What the agent is actually up against

Every JobBench task is a small dossier. Pick one role to see the details.

Role

Reporter โ€” Connecticut investigative desk

Automation desire
4.00/5
Lead in Connecticut drinking water. The state says zero water hazards. The FOIA data says otherwise.
6 sourcesยท 4 typesยท3 contradictions
Source flow
Heterogeneous inputs

Multiple Hartford-area systems exceed the 15 ppb federal action level.

conflicts with CT_2024_Surveillance_Report โ€” FOIA exceedances vs. 0% home-hazard finding
conflicts with EPA_LCRI_Factsheet โ€” Rule finalized vs. current enforcement cycle

0% of investigated homes identified water as a lead hazard.

conflicts with FOIA_water_data โ€” FOIA exceedances vs. 0% home-hazard finding

CT rows only for 2017โ€“2019; 2020โ€“2022 are dagger-marked non-submissions.

conflicts with martinez_interview โ€” CDC n=1,666 vs. Martinez 30% clinic-specific

10 ppb action level finalized Oct 2024 โ€” not yet enforceable.

conflicts with FOIA_water_data โ€” Rule finalized vs. current enforcement cycle

Pediatric referrals up 30% post-threshold change (Dr. Martinez).

conflicts with CDC_2017_2022_Blood_Lead โ€” CDC n=1,666 vs. Martinez 30% clinic-specific

Waterbury 16.1 ppb vs. Newark 47.9 ppb โ€” trajectory, not point-in-time.

Agent
reasoning over reporter sources
Deliverables
  • Thesis-driven pitch memo
  • 3-sheet data workbook
  • 15+ entry source verification log

Reasoning challenges by design

click for full detail
ยง 06 โ€” Breakdown

Heatmap

Harness: OpenCode v1.14.18 for all models; judge: Grok 4.3. โ€œโ€“โ€ = the run produced no output there โ€” the agent hit the per-task time limit, or the model refused (Claude Fable 5 cells otherwise include its Opus 4.8 fallback on refusals).

Scale0โ€“10%10โ€“20%20โ€“30%30โ€“40%40%+
Occupation
Fable557.4
MuseSpark 1.154.7
KimiK354.3
Opus4.848.4
GPT-5.6SOL45.4
Opus4.744.5
GLM5.243.4
GPT-5.538.3
Sonnet4.636.6
Gemini3.5 Flash31.5
Sonnet4.520.7
Gemini3.1 Pro15.9
Business / Financial Ops
Bookkeeping & Accounting Clerks775126โ€“53260243227617
HR Specialists6994450383875289900
Licensing Examiners / Inspectors6778727269538150672588
Management Analysts613938291744293326936
Personal Financial Advisors313121314667810183100
Purchasing Agents65615649484735355431207
Training & Development Specialists47615754544949622626317
Avg.59474547464640353323107
Office / Admin Support
Court Clerksโ€“5855โ€“29372426421002918
Customer Service Reps665350535063294229880
Data Entry Keyers898867705766725028593849
Medical Secretaries54545454464141315441158
Police / Fire Dispatchers683636573619572847173628
Secretaries & Admin Assistants52393738274831272353110
Avg.665550544146423437382619
Computer / Mathematical
Biostatisticians23575123374346442549520
CS Researchers3543372029241824179255
Statisticians435044505340534955483123
User Support Specialists65535755384861371539944
Web Administrators367248366060723636242412
Avg.415547374443503829341921
Architecture / Engineering
Civil Engineers615358415049504445323131
Mechanical Eng. Technicians61477045483936353029203
Mechanical Engineers554567555836451803390
Petroleum Engineers5256525232322832321200
Avg.57506248473940322727159
Management
Financial Managers677690555239332850432924
Health Services Managers45393338213220143416814
IT / IS Managers465448285443274330251710
Supply Chain Managers295035538518225555
Avg.475551324230252730221413
Arts / Media
Producers10064781008356787847531922
Reporters & Correspondents735763736063634723331020
Technical Writers656668594553624456435517
Avg.796270776357685642432820
Other (Legal ยท Sales ยท Science ยท Edu.)
Lawyers757575637563383850255013
Online Merchants67607769705059506035619
Securities Sales Agents2727413835145127240140
Soc. Sci. Research Assistants747175797064737074664642
Sociology Teachers (Postsec.)516352604254353639401315
Tech & Sci. Sales Reps73396333334435251419811
Avg.615664575448494144312317

Cite

@misc{li2026jobbenchaligningagentwork,
  title         = {JobBench: Aligning Agent Work With Human Will},
  author        = {Yuetai Li and Yichen Feng and Zhangchen Xu and Zixian Ma and others},
  year          = {2026},
  eprint        = {2605.26329},
  archivePrefix = {arXiv},
  primaryClass  = {cs.AI},
  url           = {https://arxiv.org/abs/2605.26329}
}