Mario Marcolongo
Model behavior testing · adversarial QA · evidence-bound reporting · evaluation operations
AI evaluation and research-verification specialist with self-directed model-behavior testing across chat, image, agentic tool-use and indirect prompt-injection challenges. The Gray Swan Proving Ground profile displayed rank #75 (top 6%) and 110 total breaks on 26 July 2026. Brings eight years of auditable claim verification and 59+ published-project experience inside a 267K-subscriber, 36.5M-view science-communication environment.
Selected evidence
Relevant experience
Model-Behavior Evaluator
- Conduct self-directed testing of LLM instruction handling, policy boundaries and edge cases across chat, image, agentic tool-use and indirect prompt-injection settings.
- Reached #75 on the Proving Ground (top 6%) with 110 platform-displayed total breaks in the 26 July 2026 snapshot.
- Document the visible 109/110 discrepancy and separate platform-reported outcomes from independent verification, model-wide conclusions or security certification.
Scientific Content Quality & Operations Contractor
- Supported 59+ documented projects within a science-communication brand with 267K YouTube subscribers and 36.5M views.
- Conduct recurring primary-literature review and scientific fact-checking; contribute assignment-specific scripts, data visualizations, slides/on-screen assets and selected performance-aware thumbnail or visual-packaging work.
- Formally acknowledged in Giacomo Moro Mauretto's Mondadori book Italiani veri for scientific-literature research and error detection.