the workforce labor index: a transaction-anchored, iosco-aligned benchmark for the commodifiable outputs of artificial labor
we introduce the workforce labor index (wli), the first transaction-anchored, iosco-aligned benchmark for the price of artificial-labor outputs. existing references for ai work are either capability leaderboards without prices (helm, chatbot arena, swe-bench, Ï„-bench, gaia) or vendor list-price pages without quality conditioning. neither is a price-discovery instrument. wli closes the structural slot held historically by sofr for short-term funding, kelley blue book for used vehicles, and bls oews for human wages.
the headline unit is the per-transaction adjusted quality output (aqo), aqoj = cj / (eÏ€,c · sj). the category headline (haqo) is a volume-weighted median over a rolling 90-day window with an 80% bootstrap confidence interval (b = 10,000 resamples). provider-level estimates (paqo) use empirical-bayes shrinkage with κ = 30, doubling as a wash-trade defense. a four-tier source hierarchy (a–d) categorically excludes list prices from the headline tier. ten sealed, sha256-receipted eval banks at v1.0 (50 deterministic tasks each) anchor eÏ€,c. we map the methodology to all 19 iosco principles (pd415, 2013) and publish a self-attested statement of compliance. v1.0 is released pre-first-transaction (tx1); every empirical headline is honestly flagged preview under all-tier-c data. methodology, calculator, sealed banks, and audit trail are released under cc-by-4.0 / apache-2.0 / cc-by-nc-4.0 for reproduction and extension.
1.introduction
the rapid commodification of artificial-labor outputs has outpaced the development of a citable, third-party benchmark for their market price. existing references are either vendor-supplied — and therefore not independent — or based on small-sample survey methods[1]. neither survives the scrutiny applied to financial benchmarks in the post-libor era.
we present the workforce labor index: a transaction-anchored consensus price for the commodifiable outputs of artificial labor, calculated weekly under an iosco-aligned methodology. the index is designed to function in the same citation chain as sofr[2] and the bloomberg short-term bank yield index[3].
1.1 contributions
this document makes four contributions: (i) an iosco-aligned methodology for ai-labor benchmarks; (ii) a published, versioned weekly index; (iii) an open task-bank protocol — the workforce eval — that admits transactions to the corpus; (iv) a reproducibility standard with an independent weekly audit.
1.2 foundations
the institution that publishes this index is itself organized under the interpretable context methodology (icm), in which folder structure is the agentic architecture[8]. the full icm paper — van clief & mcdermott (2026) — is the architectural substrate beneath the editorial, eval, and audit machinery described here, and is available in full below.
2.data
the wli corpus is built from transactions captured through the workforce marketplace and from partner vendors that submit signed transaction records under the spec at /spec/v1.0. table 1 summarizes the active corpus by category using the current illustrative sample.
2.1 admission criteria
a candidate transaction is admitted as tier a only after all of: (i) the signature validates; (ii) all required fields are present and conformant to /spec/v1.0; (iii) the buyer-confirmation marker is present and non-revoked; (iv) the output passes three-reviewer scoring with cohen’s κ ≥ 0.85.
| category | n · 7d | wli | 95% ci |
|---|
3.methodology
the wli closing figure for a category in week w is the trimmed mean of tier-a-admitted prices over the rolling 7-day window ending at week-close, with 5% winsorization at both tails. let Pw be the set of admitted prices in week w:
the per-transaction unit aggregated here — the adjusted quality-outcome (aqo), cost divided by eval × outcome — is specified in full in the companion aqo methodology paper.
3.1 confidence intervals
the bootstrap confidence interval is computed across 10,000 resamples drawn with replacement from Pw. the ci half-width is published alongside each weekly figure — a figure without an interval is a claim, not a measurement.
3.2 n-suppression
if |Pw| < 50, the figure is suppressed and a suppression flag is published in place of the value. this prevents low-power weeks from generating publishable noise.
4.sealed eval banks
the provider eval score eÏ€,c is the most-attackable input in equation (1): a vendor that controls its own eval can move haqo arbitrarily. wli’s defense is sealed, sha256-receipted, deterministic-grading task banks. for each category, a v1.0 bank comprises tasks.json (50 records with structured verification_rule — no llm judge), pass-criteria.md, and sealing-receipt.txt (sha256 of both files + seal timestamp). once sealed: true, the directory is immutable. any change requires a new sibling version folder and an entry in corrections.md — the same versioning discipline as swe-bench verified.
4.1 the v1.0 banks
at v1.0, ten sealed banks (500 deterministic tasks total) are deployed at strategy/eval-banks/<category>/v1.0/. independent verification by sha256sum (posix) or Get-FileHash (powershell) reproduces the published hash byte-for-byte.
| category | sha256 (tasks.json · first 16) | sealed (utc) |
|---|---|---|
| cs-resolution | fc5536079348cd46… | 2026-06-19 10:22z |
| sdr-bdr | 44bc635cf3f1a461… | 2026-06-19 10:10z |
| claim-adjudication | 162aabcb89b0d0ec… | 2026-06-19 11:08z |
| content-moderation | 530d8c980edf7c1b… | 2026-06-19 11:08z |
| data-extraction | e40dc02ee180bf9d… | 2026-06-19 11:00z |
| document-review | cdb83e7415a330d9… | 2026-06-19 11:38z |
| email-composition | 24855ce7a9c5f81f… | 2026-06-19 11:38z |
| lead-qualification | b4811b33f5fc3a7b… | 2026-06-19 11:00z |
| ticket-processing | 0eeb08d3e7b67091… | 2026-06-19 11:08z |
5.iosco pd415 alignment
wli v1.0 controls map to all 19 iosco principles for financial benchmarks (iosco pd415, july 2013). the full clause-by-clause statement is published as the statement of compliance v1.0. table 3 summarizes by pillar; for the full 19-principle mapping see appendix b of the arxiv preprint.
| pillar | principles | wli implementation |
|---|---|---|
| governance | p · 1–5 | Griffain sole administrator · oversight committee target 2027-06-04 |
| benchmark quality | p · 6–10 | four-tier hierarchy · 80% bootstrap ci · per-category status files |
| methodology quality | p · 11–15 | this paper + correction protocol + 30-day comment window |
| accountability | p · 16–19 | complaints procedure · external audit target q1 2027 · 10-year audit-trail retention |
6.reproducibility
reproducibility is operationalized as: an independent third party, given (a) the tier-a corpus, (b) the methodology version, (c) the sealed bank set, and (d) the equation set in §3, must reproduce the published haqo within one ci half-width and any eπ,c to four decimal places. the reference calculator (python, mit) ships at strategy/methodology/calculator/. the synthetic appendix-a dataset must reproduce haqo = 1.111 before the calculator is used on real data.
every published number passes the 3-pass pressure test: pass 1 (formula compute, reject if e·s < 0.05 or aqo > 10× prior); pass 2 (perturbation — drop highest-volume row and ±5% noise resample, accept only if perturbed ci overlaps original ≥ 70%); pass 3 (hand-validate five random rows against the primary source). live categories re-run all three passes every 30 days.
7.governance, correction protocol & availability
the wli editorial board governs the methodology. board composition and conflict-of-interest disclosures are public. methodology changes require a 30-day public comment window and a board vote; emergency changes publish immediately with retroactive comment. version increments produce a new citation; prior versions remain accessible.
a correction triggers when a misclassified row, formula error, category-boundary misapplication, constant change, or external methodology issue is identified. the affected figure stays visible with an “under correction” banner; old and new are published side-by-side for 30 days; silent retraction is forbidden. methodology and tooling are released under cc-by-4.0 (text), apache-2.0 (reference implementation), and cc-by-nc-4.0 (sealed banks). the index, methodology spec, iosco statement, reproducibility protocol, and citation tools are published at workforcebygriffain.com/methodology.