OPEN METHODS / RAW DATA / BOUNDED CLAIMS

Research agents can inspect and challenge.

Working Machines publishes methods, dated evidence, raw sanitized results, failures, and limitations. A conclusion belongs here only when its inputs and scoring can be reviewed.

DATASET / SEPTEMBER 3, 2026

First action-discovery snapshot

Four real MCP discovery calls, ranked action IDs, manually reviewed minimal sets, raw JSON, one explicit top-one miss, and client-observed timings.

METHODOLOGY / VERSION 0.1

Agent action discovery method

Define discovery recall, selection precision, schema context cost, authorization clarity, time to an executable plan, and verification completeness.

EXECUTION EVIDENCE / THREE SANITIZED READS

Verified agent execution reports

Inspect action inputs, authorization boundaries, structured outcomes, repeated verification, observed time, and explicit limits.

Publication standard

Published reports identify the dated client and protocol, prompt or action input, expected result, safe redactions, failure accounting, observed timing, and limitations. Snapshots are not presented as uptime, reliability, or customer-performance claims.

Reproducible inputs

Every report identifies the prompt or action input, client, protocol, catalog boundary, and expected action set needed to review the result.

Inspectable evidence

Sanitized traces preserve action selection, safe arguments, status, timing, failures, and verification without publishing secrets or private account data.

Bounded conclusions

Results describe only the tested configuration. Limitations, exclusions, corrections, and review dates remain visible so agents do not generalize beyond evidence.