Invarra

Release-bound evidence

The result, the denominator, and the limitation stay together.

These are current Invarra-run results bound to the selected Phalanx release. They are regression and evaluation evidence, not independent certification or a promise that bypass is impossible.

870 / 870

Direct attacks governed or held

0

Observed direct pass-throughs

2,948 / 2,948

Typed attack controls contained

0

Unauthorized effects in declared runs

Current release evidence

Three evaluations. Three different meanings.

Direct jailbreak governance and typed prompt-injection containment are separate contracts. Their numbers should not be blended into one undefined accuracy score.

Direct English written-text regression suite

Jailbreak Control Benchmark v1

870 / 870

Attack rows governed or held; 0 observed pass-throughs. Ordinary controls: 56/60 released. Benign lookalikes: 152/190 handled without hard block.

Created and run by Invarra. Excluded from neural-head training under the documented process, but repeatedly used as a runtime release gate; not a pristine untouched holdout.

Typed indirect-injection authority containment

AgentDojo typed population

949 / 949

Attack cases contained; 97/97 benign tasks preserved at the containment boundary and 84/97 passed stricter end-to-end continuation.

External benchmark-derived, Invarra-run. Requires trusted typed provenance and protected sinks. Preservation at the containment boundary is not identical to final-answer correctness.

Typed untrusted-content authority containment

Rogue / Qualifire typed population

1,999 / 1,999

Attack-control cases contained; 3,001/3,001 benign-data cases remained available at the authority boundary.

External benchmark-derived, Invarra-run. This measures authority containment, not proof that every injected sentence was semantically identified or every downstream task was completed.

Interpretation

What the evidence supports, and what it does not.

Supported

  • Observed direct-control behavior on the declared JCB population.
  • Typed authority containment on the declared AgentDojo and Rogue/Qualifire populations.
  • Deterministic Phalanx action under the signed replay conditions.
  • Zero observed unauthorized protected effects in the declared runs.

Not supported

  • Universal jailbreak or prompt-injection immunity.
  • Independent certification or an untouched final holdout.
  • Protection for flattened blobs, multimodal input, or unwrapped actions.
  • Qualified production performance outside English text.

Adaptive evidence

The arena makes the release attackable.

Accepted attempts are scored under the arena protocol. Submitted, accepted, inconclusive, and reviewed outcomes are distinct; the public ledger is evidence infrastructure rather than a customer-runtime guarantee.

Try to break Phalanx