Reproducible inputs
Every report identifies the prompt or action input, client, protocol, catalog boundary, and expected action set needed to review the result.
OPEN METHODS / RAW DATA / BOUNDED CLAIMS
Working Machines publishes methods, dated evidence, raw sanitized results, failures, and limitations. A conclusion belongs here only when its inputs and scoring can be reviewed.
Four real MCP discovery calls, ranked action IDs, manually reviewed minimal sets, raw JSON, one explicit top-one miss, and client-observed timings.
METHODOLOGY / VERSION 0.1Define discovery recall, selection precision, schema context cost, authorization clarity, time to an executable plan, and verification completeness.
EXECUTION EVIDENCE / THREE SANITIZED READSInspect action inputs, authorization boundaries, structured outcomes, repeated verification, observed time, and explicit limits.
Published reports identify the dated client and protocol, prompt or action input, expected result, safe redactions, failure accounting, observed timing, and limitations. Snapshots are not presented as uptime, reliability, or customer-performance claims.
Every report identifies the prompt or action input, client, protocol, catalog boundary, and expected action set needed to review the result.
Sanitized traces preserve action selection, safe arguments, status, timing, failures, and verification without publishing secrets or private account data.
Results describe only the tested configuration. Limitations, exclusions, corrections, and review dates remain visible so agents do not generalize beyond evidence.