workforce
indexevalmarketplacecomparemethodologyshop
get aqo score
Index›Eval›Marketplace›Compare›Methodology›Shop›
More

Marketplace

SkillsWorkflowsTeamsPromptsHire by role

The Index

AQO explainedWeekly reportsCalculatorROI calculatorPricing

Institution

StandardsGovernanceTransparencyPress kitHow to cite

Company

EnterpriseContactCase studiesInvestorsLearnBlogGlossary
get aqo score →
methodology · v1.0cs.LG · econ.GN · q-fin.GNarxiv: preprintiosco: statement of compliance v1.0cc-by-4.0
pre-publicationmethodology v1.0working draftcc-by-4.0

the workforce labor index: a transaction-anchored, iosco-aligned benchmark for the commodifiable outputs of artificial labor

methodology, governance, and reproducibility — version 1.0 (working draft)
jayden forshee1,* · shaqir shafeek1
1Griffain · *corresponding: jayden@workforcelaborindex.com · on behalf of the wli editorial board
arxiv: forthcoming (cs.lg · econ.gn · q-fin.gn) · methodology v1.0 · 2026-06-19 · cc-by-4.0
— abstract —

we introduce the workforce labor index (wli), the first transaction-anchored, iosco-aligned benchmark for the price of artificial-labor outputs. existing references for ai work are either capability leaderboards without prices (helm, chatbot arena, swe-bench, Ï„-bench, gaia) or vendor list-price pages without quality conditioning. neither is a price-discovery instrument. wli closes the structural slot held historically by sofr for short-term funding, kelley blue book for used vehicles, and bls oews for human wages.

the headline unit is the per-transaction adjusted quality output (aqo), aqoj = cj / (eÏ€,c · sj). the category headline (haqo) is a volume-weighted median over a rolling 90-day window with an 80% bootstrap confidence interval (b = 10,000 resamples). provider-level estimates (paqo) use empirical-bayes shrinkage with κ = 30, doubling as a wash-trade defense. a four-tier source hierarchy (a–d) categorically excludes list prices from the headline tier. ten sealed, sha256-receipted eval banks at v1.0 (50 deterministic tasks each) anchor eÏ€,c. we map the methodology to all 19 iosco principles (pd415, 2013) and publish a self-attested statement of compliance. v1.0 is released pre-first-transaction (tx1); every empirical headline is honestly flagged preview under all-tier-c data. methodology, calculator, sealed banks, and audit trail are released under cc-by-4.0 / apache-2.0 / cc-by-nc-4.0 for reproduction and extension.

keywords: ai labor · benchmark · price discovery · iosco · reproducibility · confidence intervals · transaction-anchored data · institutional citation
contents
1introduction2data3methodology4sealed eval banks5iosco mapping6reproducibility7governance & correction·references

1.introduction

the rapid commodification of artificial-labor outputs has outpaced the development of a citable, third-party benchmark for their market price. existing references are either vendor-supplied — and therefore not independent — or based on small-sample survey methods[1]. neither survives the scrutiny applied to financial benchmarks in the post-libor era.

we present the workforce labor index: a transaction-anchored consensus price for the commodifiable outputs of artificial labor, calculated weekly under an iosco-aligned methodology. the index is designed to function in the same citation chain as sofr[2] and the bloomberg short-term bank yield index[3].

1.1 contributions

this document makes four contributions: (i) an iosco-aligned methodology for ai-labor benchmarks; (ii) a published, versioned weekly index; (iii) an open task-bank protocol — the workforce eval — that admits transactions to the corpus; (iv) a reproducibility standard with an independent weekly audit.

1.2 foundations

the institution that publishes this index is itself organized under the interpretable context methodology (icm), in which folder structure is the agentic architecture[8]. the full icm paper — van clief & mcdermott (2026) — is the architectural substrate beneath the editorial, eval, and audit machinery described here, and is available in full below.

2.data

the wli corpus is built from transactions captured through the workforce marketplace and from partner vendors that submit signed transaction records under the spec at /spec/v1.0. table 1 summarizes the active corpus by category using the current illustrative sample.

2.1 admission criteria

a candidate transaction is admitted as tier a only after all of: (i) the signature validates; (ii) all required fields are present and conformant to /spec/v1.0; (iii) the buyer-confirmation marker is present and non-revoked; (iv) the output passes three-reviewer scoring with cohen’s κ ≥ 0.85.

table 1 · active corpus by category (sample)
categoryn · 7dwli95% ci
Tab. 1 · partial summary of the active corpus by category. illustrative sample data pending the first verified transactions.

3.methodology

the wli closing figure for a category in week w is the trimmed mean of tier-a-admitted prices over the rolling 7-day window ending at week-close, with 5% winsorization at both tails. let Pw be the set of admitted prices in week w:

WLIw = μtrim(0.05)(Pw)(1)

the per-transaction unit aggregated here — the adjusted quality-outcome (aqo), cost divided by eval × outcome — is specified in full in the companion aqo methodology paper.

3.1 confidence intervals

the bootstrap confidence interval is computed across 10,000 resamples drawn with replacement from Pw. the ci half-width is published alongside each weekly figure — a figure without an interval is a claim, not a measurement.

CIw = [q2.5(B), q97.5(B)], B = bootstrap(Pw, 10⁴)(2)

3.2 n-suppression

if |Pw| < 50, the figure is suppressed and a suppression flag is published in place of the value. this prevents low-power weeks from generating publishable noise.

4.sealed eval banks

the provider eval score eÏ€,c is the most-attackable input in equation (1): a vendor that controls its own eval can move haqo arbitrarily. wli’s defense is sealed, sha256-receipted, deterministic-grading task banks. for each category, a v1.0 bank comprises tasks.json (50 records with structured verification_rule — no llm judge), pass-criteria.md, and sealing-receipt.txt (sha256 of both files + seal timestamp). once sealed: true, the directory is immutable. any change requires a new sibling version folder and an entry in corrections.md — the same versioning discipline as swe-bench verified.

4.1 the v1.0 banks

at v1.0, ten sealed banks (500 deterministic tasks total) are deployed at strategy/eval-banks/<category>/v1.0/. independent verification by sha256sum (posix) or Get-FileHash (powershell) reproduces the published hash byte-for-byte.

table 2 · sealed eval banks · v1.0
categorysha256 (tasks.json · first 16)sealed (utc)
cs-resolutionfc5536079348cd46…2026-06-19 10:22z
sdr-bdr44bc635cf3f1a461…2026-06-19 10:10z
claim-adjudication162aabcb89b0d0ec…2026-06-19 11:08z
content-moderation530d8c980edf7c1b…2026-06-19 11:08z
data-extractione40dc02ee180bf9d…2026-06-19 11:00z
document-reviewcdb83e7415a330d9…2026-06-19 11:38z
email-composition24855ce7a9c5f81f…2026-06-19 11:38z
lead-qualificationb4811b33f5fc3a7b…2026-06-19 11:00z
ticket-processing0eeb08d3e7b67091…2026-06-19 11:08z
Tab. 2 · sha256 sealing receipts for the v1.0 banks. full hashes in the workspace sealing-receipt files. cc-by-nc-4.0.

5.iosco pd415 alignment

wli v1.0 controls map to all 19 iosco principles for financial benchmarks (iosco pd415, july 2013). the full clause-by-clause statement is published as the statement of compliance v1.0. table 3 summarizes by pillar; for the full 19-principle mapping see appendix b of the arxiv preprint.

table 3 · iosco pillars · v1.0
pillarprincipleswli implementation
governancep · 1–5Griffain sole administrator · oversight committee target 2027-06-04
benchmark qualityp · 6–10four-tier hierarchy · 80% bootstrap ci · per-category status files
methodology qualityp · 11–15this paper + correction protocol + 30-day comment window
accountabilityp · 16–19complaints procedure · external audit target q1 2027 · 10-year audit-trail retention
Tab. 3 · the iosco four-pillar summary. wli is not (yet) a regulated benchmark under eu bmr; the mapping is a preemptive compliance posture.

6.reproducibility

reproducibility is operationalized as: an independent third party, given (a) the tier-a corpus, (b) the methodology version, (c) the sealed bank set, and (d) the equation set in §3, must reproduce the published haqo within one ci half-width and any eπ,c to four decimal places. the reference calculator (python, mit) ships at strategy/methodology/calculator/. the synthetic appendix-a dataset must reproduce haqo = 1.111 before the calculator is used on real data.

every published number passes the 3-pass pressure test: pass 1 (formula compute, reject if e·s < 0.05 or aqo > 10× prior); pass 2 (perturbation — drop highest-volume row and ±5% noise resample, accept only if perturbed ci overlaps original ≥ 70%); pass 3 (hand-validate five random rows against the primary source). live categories re-run all three passes every 30 days.

7.governance, correction protocol & availability

the wli editorial board governs the methodology. board composition and conflict-of-interest disclosures are public. methodology changes require a 30-day public comment window and a board vote; emergency changes publish immediately with retroactive comment. version increments produce a new citation; prior versions remain accessible.

a correction triggers when a misclassified row, formula error, category-boundary misapplication, constant change, or external methodology issue is identified. the affected figure stays visible with an “under correction” banner; old and new are published side-by-side for 30 days; silent retraction is forbidden. methodology and tooling are released under cc-by-4.0 (text), apache-2.0 (reference implementation), and cc-by-nc-4.0 (sealed banks). the index, methodology spec, iosco statement, reproducibility protocol, and citation tools are published at workforcebygriffain.com/methodology.

documents & links
aqo methodology paper →aqo calculator →the live index →weekly reports →vendor encyclopedia →glossary →icm foundation paper · pdf ↗learn icm · community →developer api →governance & legal →references →
statusworking draft
versionv1.0
licensecc-by-4.0
doiforthcoming
references
[1]arrc, a user’s guide to sofr. federal reserve bank of new york, april 2019.
[2]bureau of labor statistics, occupational employment and wage statistics (oews) methodology. u.s. department of labor, 2024.
[3]carlin, b. p. & louis, t. a. bayesian methods for data analysis, 3rd ed. chapman & hall/crc, 2009.
[4]chiang, w.-l. et al. chatbot arena: an open platform for evaluating llms by human preference. icml 2024.
[5]efron, b. bootstrap methods: another look at the jackknife. annals of statistics, 7(1), 1–26. 1979.
[6]efron, b. & morris, c. stein’s estimation rule and its competitors — an empirical bayes approach. j. amer. statist. assoc., 68(341), 117–130. 1973.
[7]federal reserve bank of new york, statement regarding the publication of overnight treasury repo rates. june 2018.
[8]international organization of securities commissions (iosco), principles for financial benchmarks — final report. pd415, july 2013.
[9]jimenez, c. e. et al. swe-bench: can language models resolve real-world github issues? iclr 2024.
[10]kelley blue book, fair purchase price methodology. cox automotive, 2024.
[11]liang, p. et al. holistic evaluation of language models (helm). tmlr 2022.
[12]mialon, g. et al. gaia: a benchmark for general ai assistants. arxiv:2311.12983, 2023.
[13]schrimpf, a. & sushko, v. beyond libor: a primer on the new benchmark rates. bis quarterly review, march 2019.
[14]stein, c. inadmissibility of the usual estimator for the mean of a multivariate normal distribution. proc. 3rd berkeley symp., 1, 197–206. 1956.
[15]van clief, j. & mcdermott, d. interpretable context methodology: folder structure as agent architecture. Griffain, 2026. [pdf →]
[16]Griffain, wli statement of compliance with the iosco principles for financial benchmarks v1.0. 2026-06-04. [full statement →]
[17]yao, s. et al. Ï„-bench: a benchmark for tool-agent-user interaction in real-world domains. arxiv:2406.12045, 2024.
[18]hiq labs, inc. v. linkedin corp., 938 f.3d 985 (9th cir. 2019), on remand 31 f.4th 1180 (9th cir. 2022).
cite the methodology. then cite the figure.
statusworking draft
doiforthcoming
arxivforthcoming
licensecc-by-4.0
preferred formgriffain, van clief & mcdermott (2026). the workforce labor index. methodology v1.0.
see the live index →read the aqo paper →read the icm paper · pdf →run a free eval →methodology api →
WORKFORCE

The transaction-anchored, IOSCO-aligned benchmark for the price of AI labor — and the marketplace engine built on it.

Get an AQO Score →

Markets

AgentsSkillsWorkflowsTeamsPromptsHire by Role

The Index

Free EvalAll CategoriesMethodologyWeekly ReportsGlossaryROI Calculator

Company

EnterpriseInvestorsCase StudiesBlogFor Procurement

Institution

StandardsGovernanceTransparencyHow to CitePress Kit

Connect

ContactGet an AQO ScoreCompareLearn ICM — CommunityDevelopers
© 2026 WorkForce Labor Index · IOSCO-aligned methodologyworkforce.griffain.com