Invarra

Historical Phalanx evaluations

This archive preserves evaluations of an earlier Phalanx runtime. Its jailbreak and typed-containment results do not measure the Phalanx V1 execution gateway. Statements within the archived material describe the runtime and publication context at that time.

Publication snapshot archived 19 September 2026

870 / 870

Direct attacks governed or held

0

Observed direct pass-throughs

2,948 / 2,948

Typed attack controls contained

0

Unauthorized effects in declared runs

Current release evidence

Three evaluations. Three different meanings.

Direct jailbreak governance and typed prompt-injection containment are separate contracts. Their numbers should not be blended into one undefined accuracy score.

Direct written-text regression suite

Jailbreak Control Benchmark v1

870 / 870

Attack rows governed or held; 0 observed pass-throughs. Ordinary controls: 56/60 released. Benign lookalikes: 154/190 handled without hard block.

Created and run by Invarra. Excluded from neural-head training under the documented process, used as a runtime release gate.

Typed indirect-injection authority containment

AgentDojo typed population

949 / 949

Attack cases contained; 97/97 benign tasks preserved at the containment boundary and 84/97 passed stricter end-to-end continuation.

External benchmark-derived, Invarra-run. Requires trusted typed provenance and protected sinks. Preservation at the containment boundary is not identical to final-answer correctness.

Typed untrusted-content authority containment

Rogue / Qualifire typed population

1,999 / 1,999

Attack-control cases contained; 3,001/3,001 benign-data cases remained available at the authority boundary.

External benchmark-derived, Invarra-run. This measures authority containment, not proof that every injected sentence was semantically identified or every downstream task was completed.

JCB is a maintained release regression suite, not a pristine untouched holdout.

Interpretation

What the evidence supports.

Supported

  • Observed direct-control behavior on the declared JCB population.
  • Typed authority containment on the declared AgentDojo and Rogue/Qualifire populations.
  • Deterministic Phalanx action under the signed replay conditions.
  • Zero observed unauthorized protected effects in the declared runs.