AI evaluation and model-behavior record.
Self-directed adversarial testing across chat, image, agentic tool-use and indirect prompt-injection challenges. This page explains what the public activity demonstrates, how the work is approached and where the evidence stops.
What the record demonstrates
Sustained exploratory testing, familiarity with several model-interaction surfaces, repeated submission workflows and careful distinction between public evidence and interpretation.
Instruction hierarchy and multi-turn behavior
Probe whether constraints, refusals and contextual priorities remain stable across direct and extended conversations.
Multimodal and visual inputs
Test interactions between textual instructions and information embedded in images, diagrams or visual typography.
Agentic tool-use and function selection
Investigate whether tool selection and function calls remain aligned when external, nested or conflicting context is present.
Indirect prompt injection
Test how systems handle adversarial instructions contained in retrieved documents, web pages or uploaded files.
Evaluation approach
The current public record is stronger as evidence of evaluation judgment and persistence than as evidence of security engineering. The methodology therefore emphasizes observable behavior, test variation, documentation and evidence boundaries.
Define the behavior under test
Translate an ambiguous risk or product behavior into an observable condition, boundary or expected response.
Vary context and attack path
Explore direct, multi-turn, multimodal, nested and tool-mediated variants rather than relying on a single prompt.
Record evidence conservatively
Separate observed behavior and platform status from interpretation, model-wide conclusions or unsupported severity claims.
Communicate reproducibly
Document the test area, preconditions, behavior, limitations and available public evidence in a form another evaluator can inspect.
Publicly inspectable evidence
The live Gray Swan profile is the primary source. The latest dated evidence page records the later 25 July metrics, the original screenshot hash and capture metadata, while keeping Proving Ground and Arena figures separate. It also records that the four visible area counters sum to 109 rather than silently inferring an aggregation rule.
26 July 2026 Gray Swan profile screenshot
Dated public profile record for #75, top 6%, 110 Proving Ground total breaks, Arena rank #371, 27 global unique breaks, 1,090 points and 246 submissions.
Open the dated evidence and preservation record → · Previous 25 July snapshot → · Previous evidence route →
damage-property
Category visible through the public platform record. It is not presented as a complete vulnerability report or independently reproduced exploit.
toxic-plant
Category visible through the public platform record. It is not presented as a complete vulnerability report or independently reproduced exploit.
package-theft-image
Category visible through the public platform record. It is not presented as a complete vulnerability report or independently reproduced exploit.
Open the historical 24 July 26-wave activity table
| Wave | Leaderboard-counted breaks | Available challenges shown |
|---|---|---|
| Wave 1 | 14 | 67 |
| Wave 2 | 10 | 72 |
| Wave 3 | 10 | 72 |
| Wave 4 | 2 | 46 |
| Wave 5 | 6 | 72 |
| Wave 6 | 10 | 72 |
| Wave 7 | 2 | 64 |
| Wave 8 | 0 | 67 |
| Wave 9 | 0 | 67 |
| Wave 10 | 3 | 67 |
| Wave 11 | 2 | 67 |
| Wave 12 | 3 | 67 |
| Wave 13 | 0 | 67 |
| Wave 14 | 0 | 67 |
| Wave 15 | 8 | 67 |
| Wave 16 | 3 | 67 |
| Wave 17 | 14 | 75 |
| Wave 18 | 0 | 75 |
| Wave 19 | 0 | 56 |
| Wave 20 | 0 | 56 |
| Wave 21 | 0 | 56 |
| Wave 22 | 6 | 56 |
| Wave 23 | 3 | 57 |
| Wave 24 | 2 | 56 |
| Wave 25 | 3 | 55 |
| Wave 26 | 4 | 56 |
Limitations and interpretation
This record supports applications to AI evaluation, safeguards operations, trust & safety and adversarial QA roles. It does not yet establish the application-security, independent Python development or automated-framework experience expected from senior AI red-team engineers.
- The 25 July screenshot displays 110 Total Breaks while its four visible category counters sum to 109; both values are reported exactly and no internal aggregation rule is inferred.
- The historical 24 July record remains available separately and documented a 105/106 discrepancy before the later profile update.
- The public source does not expose complete prompts, outputs, model versions or adjudication materials for every result.
- Counts and rankings are snapshots and may change after 26 July 2026.
- No claim is made that every recorded result represents a security vulnerability or model-wide failure.
- The next evidence milestone is a public evaluation dataset with explicit rubrics, reproducible test cases and calibration with another evaluator.