Mario Marcolongo

AI Evaluation & Model Behavior Specialist

Model behavior testing · adversarial QA · evidence-bound reporting · evaluation operations

Italy · Italian/EU citizen · Open to worldwide relocation · International B2B contractingmariomarcolongo.com[email protected]LinkedInGitHub

AI evaluation and research-verification specialist with self-directed model-behavior testing across chat, image, agentic tool-use and indirect prompt-injection challenges. The Gray Swan Proving Ground profile displayed rank #74 (top 6%) and 113 total breaks on 29 July 2026. Supporting work demonstrates multilingual content quality, consumer-genomics privacy research, corporate-source reconciliation, archival source recovery, claim-to-source auditing and evidence-bound reporting across sensitive records.

Selected evidence

113Platform-displayed Proving Ground breaks#74 · top 6% · 29 July 2026
255Arena submissions shown28 global unique breaks · 1,120 points
80Published content contributions55 YouTube videos · 4 articles · 21 short-form pieces
4,317Auditable Wikimedia contributionsEight years of inspectable claim and source work

Relevant experience

Model-Behavior Evaluator

Independent practice · Gray Swan Proving Ground
Jul 2026 — Present
  • Conduct self-directed testing of LLM instruction handling, policy boundaries and edge cases across chat, image, agentic tool-use and indirect prompt-injection settings.
  • Reached #74 on the Proving Ground (top 6%) with 113 platform-displayed total breaks in the 29 July 2026 snapshot.
  • Document the visible 112/113 discrepancy and separate platform-reported outcomes from independent verification, model-wide conclusions or security certification.

Scientific Content Quality & Operations Contractor

Entropy for Life · Independent contractor
Jun 2023 — Present
  • Delivered 80 documented published content contributions: 55 YouTube videos · 4 articles · 21 short-form pieces.
  • Conduct recurring primary-literature review, scientific fact-checking and English-to-Italian localization; adapt evidence into Italian scripts, visualizations and short-form content while preserving terminology, meaning and source context.
  • Designed and built entropyforlife.it in WordPress; formally acknowledged in Giacomo Moro Mauretto's Mondadori book Italiani veri for scientific-literature research and error detection.

Additional relevant experience

Founder & Research-Workflow Owner

Yourself to Science™
Aug 2024 — Present
  • Founded and operate an open-source directory indexing more than 55 clinical studies, biobanks, donation programs, registries and other research initiatives.
  • Defined inclusion criteria, verification workflows, metadata structure, provenance requirements and licensing boundaries.
  • Define requirements, inspect code structure and behavior, test implementations and guide AI-assisted technical iteration through deployment and maintenance.

Additional evidence

Research-integrity product operations

Created and operate Notandia (formerly MDPI Filter), a browser and Zotero research-integrity tool. It helps researchers identify articles from publishers whose editorial and peer-review practices have attracted scrutiny—including MDPI and Frontiers—and checks Crossref/Retraction Watch records for formal notices such as retractions, corrections and expressions of concern. I define the evidence rules, privacy safeguards, ambiguity handling, false-positive boundaries and release tests.

Structured evaluation practice

Repeated adversarial testing across instruction hierarchy, multimodal inputs, tool-use behavior and untrusted external context.

Investigation and evidence-bound judgment

Attributed public cases demonstrate source synthesis, privacy-policy analysis, corporate-source reconciliation, archive recovery, legal-stage chronology, source-quality auditing, scientific taxonomy design and explicit separation of evidence from inference.

Capabilities

AI safety testingExploratory adversarial testing, prompt and jailbreak analysis, multi-turn behavior, multimodal inputs, agentic tool-use and indirect prompt injection
Evaluation operationsTest planning, evidence capture, reproducibility notes, taxonomy thinking, issue classification, severity-oriented reporting and mitigation-retesting concepts
OSINT and research verificationPublic-source and bibliographic research, claim decomposition, web-archive recovery, source-provenance analysis, corporate and legal record reconciliation, cross-source corroboration and evidence-bound reporting
Technical operationsCodebase reading and behavior inspection, requirements definition, functional testing, AI-assisted implementation workflows, Git/GitHub, JSON, REST APIs, WordPress, Cloudflare Pages and AWS Lambda deployment
Multilingual quality and communicationItalian-native and English-C1 content review, English-to-Italian scientific localization, terminology consistency, source-faithful adaptation and clear evidence limitations for technical and non-technical audiences

Best-fit role families

  • AI evaluation and safeguards operations
  • AI content red teaming and adversarial QA
  • Model behavior, trust & safety and policy testing
  • Human-data quality, grading and evaluation operations

Credentials & language

  • GALENOS Crowd Evidence Synthesis Training — Cochrane Crowd & GALENOS, 2026
  • Career Essentials in Generative AI — Microsoft & LinkedIn, 2024
  • EF SET English Certificate — 68/100, C1 overall, 2024
  • Italian — native. English — C1 overall (EF SET 68/100); advanced technical reading and professional writing, with practical English-to-Italian scientific localization and cross-language content-quality work.