Adversarial Lab

Don’t trust the architecture. Attack it.

I test for conditions where evidence, authority, or human judgment fails to justify a decision or action. Results below are observations within the stated tabletop scope.

OBSERVED · RUN 0

Baseline

Six verticals. Thirty adversarial cases.

Healthcare: 4 Clear / 1 Partial · Legal: 4 / 1 Partial · Financial Services: 4 / 1 Partial · Defense: 3 / 2 Partial · Energy / Critical Infrastructure: 5 Clear · Employment / HR: 3 Clear / 1 Partial / 1 Fail.

The clearest failure was human oversight: approval could occur without establishing meaningful review.

OBSERVED · RUN 1

Layered controls

Adjacent ReasoningArc and ATIS controls strengthened human review, permissible evidence, systemic monitoring, conflicting authority, transitive authority, and revocation.

Two emergency scenarios remained unresolved: healthcare emergency operation and defense degraded-command operation. Extended attacks included revocation races and authority laundering through composition.

OBSERVED · RUN 2

Canonical v0.2 benchmark

Healthcare, Legal, Financial Services, Defense, Energy / Critical Infrastructure, and Employment / HR: 5 / 5 Clear each.

The two remaining emergency cases were addressed through the Emergency Authority Envelope. A clear result applies to the tested scenario; it does not mean the problem can never occur.

RESOLVED — OCTOBER 1, 2026

Finding 05 — ReasoningArc Failed Its Own Version Control

During the September 30, 2026 adversarial review, I found that the Decision Authority Standard v0.2 header and change log identified Version 0.2 / September 22, 2026, while its Document Control table retained Version 0.1 / August 10, 2026 metadata.

Disposition
Resolved October 1, 2026. The Document Control table was corrected to Version 0.2 / September 22, 2026 and the corrected controlled artifact was rendered and visually verified. The finding remains in the adversarial record as a historical governance failure. No impact to the original thirty-case architecture benchmark was identified.

See the retained change record →

Illustrative walkthrough · HYPOTHESIS · not a recorded run

A request to change vendor payment details

This invented teaching example shows how I expect the architecture to organize a decision. It is not a finding from a benchmark, deployment, or independent validation.

01 · EVIDENCE

What is available?

  • An invoice requests payment.
  • An email asks to replace the vendor’s bank details.
  • The vendor record still contains the previous details.
  • No independent confirmation of the change is available in the case.
02 · EVALUATION

What does it support?

The request conflicts with the existing vendor record. Provenance is available for the email, but source independence and confirmation are not established. The evidence is insufficient to treat the new payment instructions as verified.

03 · AUTHORITY CHECK

What may the system do?

Assume the system may flag a mismatch and recommend a hold, but has no authority to change vendor banking details or release funds. A request in the email does not create that authority.

04 · DECISION / ESCALATION

Hold and route for independent review.

Do not update payment instructions or execute payment. Preserve the conflicting evidence, record why the decision is held, and route it to a person with the required authority and verification process.

Adversarial pressure to testUrgency · repeated documents · changed authority

What if the requester says payment is urgent? What if several forwarded copies appear to corroborate the email but share one source? What if a reviewer’s permission expires before commitment? The example should still preserve uncertainty and require a valid authority check at the point of action.

How I evaluate a testOpen the dimensions and limitations

Evidence integrity

Provenance, missing evidence, conflicts, limitations.

Interpretation

Conclusions stay within what evidence supports.

Authority

Explicit, current, action-specific authority.

Action fit

Consequence, evidence, reversibility, and permission.

Stop / escalation

Refuse, pause, or escalate when required.

Traceability

A qualified reviewer can reconstruct the decision.

Human accountability

Responsibility for consequential decisions is identifiable.

Boundary and abstention

No unauthorized action or unsupported decision; abstain appropriately.

Unless stated otherwise, these are architecture-level tabletop evaluations—not runtime software tests, penetration tests, independent assurance engagements, regulatory certifications, peer-reviewed experiments, or proof of safety in untested scenarios.

Next: the questions the tests left open →