Public evaluation evidence

AI evaluation and model-behavior record.

Self-directed adversarial testing across chat, image, agentic tool-use and indirect prompt-injection challenges. This page explains what the public activity demonstrates, how the work is approached and where the evidence stops.

110Proving Ground total breaks#75 · top 6% · 26 July 2026
246Arena submissions#371 Arena rank · 27 unique breaks · 1,090 points
#75 · Top 6%Proving Ground rankTime-sensitive platform snapshot · 26 July 2026

What the record demonstrates

Sustained exploratory testing, familiarity with several model-interaction surfaces, repeated submission workflows and careful distinction between public evidence and interpretation.

AREA 01

Instruction hierarchy and multi-turn behavior

Probe whether constraints, refusals and contextual priorities remain stable across direct and extended conversations.

AREA 02

Multimodal and visual inputs

Test interactions between textual instructions and information embedded in images, diagrams or visual typography.

AREA 03

Agentic tool-use and function selection

Investigate whether tool selection and function calls remain aligned when external, nested or conflicting context is present.

AREA 04

Indirect prompt injection

Test how systems handle adversarial instructions contained in retrieved documents, web pages or uploaded files.

Evaluation approach

The current public record is stronger as evidence of evaluation judgment and persistence than as evidence of security engineering. The methodology therefore emphasizes observable behavior, test variation, documentation and evidence boundaries.

STEP 01

Define the behavior under test

Translate an ambiguous risk or product behavior into an observable condition, boundary or expected response.

STEP 02

Vary context and attack path

Explore direct, multi-turn, multimodal, nested and tool-mediated variants rather than relying on a single prompt.

STEP 03

Record evidence conservatively

Separate observed behavior and platform status from interpretation, model-wide conclusions or unsupported severity claims.

STEP 04

Communicate reproducibly

Document the test area, preconditions, behavior, limitations and available public evidence in a form another evaluator can inspect.

Publicly inspectable evidence

The live Gray Swan profile is the primary source. The latest dated evidence page records the later 25 July metrics, the original screenshot hash and capture metadata, while keeping Proving Ground and Arena figures separate. It also records that the four visible area counters sum to 109 rather than silently inferring an aggregation rule.

PUBLIC PLATFORM LABEL

damage-property

Category visible through the public platform record. It is not presented as a complete vulnerability report or independently reproduced exploit.

PUBLIC PLATFORM LABEL

toxic-plant

Category visible through the public platform record. It is not presented as a complete vulnerability report or independently reproduced exploit.

PUBLIC PLATFORM LABEL

package-theft-image

Category visible through the public platform record. It is not presented as a complete vulnerability report or independently reproduced exploit.

Open the historical 24 July 26-wave activity table
WaveLeaderboard-counted breaksAvailable challenges shown
Wave 11467
Wave 21072
Wave 31072
Wave 4246
Wave 5672
Wave 61072
Wave 7264
Wave 8067
Wave 9067
Wave 10367
Wave 11267
Wave 12367
Wave 13067
Wave 14067
Wave 15867
Wave 16367
Wave 171475
Wave 18075
Wave 19056
Wave 20056
Wave 21056
Wave 22656
Wave 23357
Wave 24256
Wave 25355
Wave 26456

Limitations and interpretation

This record supports applications to AI evaluation, safeguards operations, trust & safety and adversarial QA roles. It does not yet establish the application-security, independent Python development or automated-framework experience expected from senior AI red-team engineers.

  • The 25 July screenshot displays 110 Total Breaks while its four visible category counters sum to 109; both values are reported exactly and no internal aggregation rule is inferred.
  • The historical 24 July record remains available separately and documented a 105/106 discrepancy before the later profile update.
  • The public source does not expose complete prompts, outputs, model versions or adjudication materials for every result.
  • Counts and rankings are snapshots and may change after 26 July 2026.
  • No claim is made that every recorded result represents a security vulnerability or model-wide failure.
  • The next evidence milestone is a public evaluation dataset with explicit rubrics, reproducible test cases and calibration with another evaluator.