Mario Marcolongo

AI Evaluation & Model Behavior Specialist

Model behavior testing · adversarial QA · evidence-bound reporting · evaluation operations

Italy · EU work authorization · Open to worldwide relocation and international B2B contractingmariomarcolongo.com[email protected]LinkedInGitHub

AI evaluation and research-verification specialist with self-directed model-behavior testing across chat, image, agentic tool-use and indirect prompt-injection challenges. The Gray Swan Proving Ground profile displayed rank #75 (top 6%) and 110 total breaks on 26 July 2026. Brings eight years of auditable claim verification and 59+ published-project experience inside a 267K-subscriber, 36.5M-view science-communication environment.

Selected evidence

110Platform-displayed Proving Ground breaks#75 · top 6% · 26 July 2026
246Arena submissions shown27 global unique breaks · 1,090 points
59+Scientific projects supported55+ YouTube projects · 4 articles
4,317Auditable Wikimedia contributionsEight years of inspectable claim and source work

Relevant experience

Model-Behavior Evaluator

Independent practice · Gray Swan Proving Ground
Jul 2026 — Present
  • Conduct self-directed testing of LLM instruction handling, policy boundaries and edge cases across chat, image, agentic tool-use and indirect prompt-injection settings.
  • Reached #75 on the Proving Ground (top 6%) with 110 platform-displayed total breaks in the 26 July 2026 snapshot.
  • Document the visible 109/110 discrepancy and separate platform-reported outcomes from independent verification, model-wide conclusions or security certification.

Scientific Content Quality & Operations Contractor

Entropy for Life · Independent contractor
Jun 2023 — Present
  • Supported 59+ documented projects within a science-communication brand with 267K YouTube subscribers and 36.5M views.
  • Conduct recurring primary-literature review and scientific fact-checking; contribute assignment-specific scripts, data visualizations, slides/on-screen assets and selected performance-aware thumbnail or visual-packaging work.
  • Formally acknowledged in Giacomo Moro Mauretto's Mondadori book Italiani veri for scientific-literature research and error detection.

Additional relevant experience

Founder & Research-Workflow Owner

Yourself to Science™
Aug 2024 — Present
  • Founded and operate an open-source directory indexing more than 55 clinical studies, biobanks, donation programs, registries and other research initiatives.
  • Defined inclusion criteria, verification workflows, metadata structure, provenance requirements and licensing boundaries.
  • Specify requirements, inspect code structure and behavior, test implementations and guide AI-assisted technical iteration; do not claim independent software development.

Additional evidence

Research-integrity product operations

Defined and verified cross-surface behavior for MDPI Filter across browser targets and Zotero, with exact-evidence matching, false-positive avoidance, privacy controls and reproducible release checks.

Structured evaluation practice

Repeated adversarial testing across instruction hierarchy, multimodal inputs, tool-use behavior and untrusted external context.

Human-subject research operations

Co-developed and co-facilitated structured remote focus groups with autistic participants on sexuality and relationships, supervised by Marta Panzeri at the University of Padua DPSS.

Capabilities

AI safety testingExploratory adversarial testing, prompt and jailbreak analysis, multi-turn behavior, multimodal inputs, agentic tool-use and indirect prompt injection
Evaluation operationsTest planning, evidence capture, reproducibility notes, taxonomy thinking, issue classification, severity-oriented reporting and mitigation-retesting concepts
Research verificationPrimary-source fact-checking, bibliographic research, claim decomposition, source-quality assessment, cross-source corroboration and evidence screening
Technical operationsCodebase reading and behavior inspection, requirements definition, functional testing, AI-assisted implementation workflows, Git/GitHub, JSON, REST APIs, WordPress, Cloudflare Pages and AWS Lambda deployment
CommunicationTechnical and professional English writing, concise findings, evidence limitations and explanations for mixed technical/non-technical audiences

Best-fit role families

  • AI evaluation and safeguards operations
  • AI content red teaming and adversarial QA
  • Model behavior, trust & safety and policy testing
  • Human-data quality, grading and evaluation operations

Credentials & language

  • GALENOS Crowd Evidence Synthesis Training — Cochrane Crowd & GALENOS, 2026
  • Career Essentials in Generative AI — Microsoft & LinkedIn, 2024
  • EF SET English Certificate — 68/100, C1 overall, 2024
  • Italian — native. English — C1 overall (EF SET 68/100); advanced technical reading and professional/technical writing.