---
title: "Mario Marcolongo — AI Evaluation &amp; Model Behavior Analyst"
description: "Two-page role-family application CV for ai evaluation &amp; model behavior analyst roles."
---

[Mario Marcolongo](/)

Role-family application CV

## AI Evaluation & Model Behavior Analyst

This is the concise, role-family document to send for this role. The master CV is a comprehensive evidence archive; it is normally not attached unless requested.

[Research & Data Quality CV](/cv-research)[All application CVs](/#applications)[Master CV](/cv)Print / Save PDF

# Mario Marcolongo

AI Evaluation & Model Behavior Analyst

Model testing · adversarial QA · evidence reporting · test planning

Italy · EU/EEA work-authorised · Open to sponsored international relocation and B2B engagements[mariomarcolongo.com](https://mariomarcolongo.com)[me@mariomarcolongo.com](mailto:me@mariomarcolongo.com)[LinkedIn](https://www.linkedin.com/in/mario-marcolongo)[GitHub](https://github.com/jnton)

AI evaluation and research-verification analyst with self-directed model-behavior testing across chat, image, agentic tool-use and indirect prompt-injection challenges. The Gray Swan Proving Ground profile displayed rank #74 (top 6%) and 113 total breaks on 29 July 2026. Supporting work demonstrates multilingual content quality, consumer-genomics privacy research, corporate-source reconciliation, archival source recovery, claim-to-source auditing and reporting that separates evidence from inference across sensitive records.

## Selected evidence

**113**Platform-displayed Proving Ground breaks#74 · top 6% · 29 July 2026

**255**Arena submissions shown28 global unique breaks · 1,120 points

**80**Published content contributions55 YouTube videos · 4 articles · 21 short-form pieces

**4,317**Auditable Wikimedia contributionsEight years of inspectable claim and source work

## Relevant experience

### Independent AI Evaluator

Independent practice · Gray Swan Proving Ground

[Evaluation record ↗](https://mariomarcolongo.com/security)[Dated evidence ↗](https://mariomarcolongo.com/evidence/gray-swan-2026-07-29/)[Public profile ↗](https://app.grayswan.ai/arena/user/6a57be70d15e123775a1e9cf)

Jul 2026 — Present

-   Conduct self-directed testing of LLM instruction handling, policy boundaries and edge cases across chat, image, agentic tool-use and indirect prompt-injection settings.
-   Reached #74 on the Proving Ground (top 6%) with 113 platform-displayed total breaks in the 29 July 2026 snapshot.
-   Document the visible 112/113 discrepancy and separate platform-reported outcomes from independent verification, model-wide conclusions or security certification.

### Science Writer & Fact-Checker / Website Manager (WordPress)

Entropy for Life · Independent contractor

[Official Entropy for Life work record ↗](https://entropyforlife.it/mario-marcolongo-entropy-for-life/)

Jun 2023 — Present

-   Delivered 80 documented published content contributions: 55 YouTube videos · 4 articles · 21 short-form pieces.
-   Conduct recurring primary-literature review, scientific fact-checking and English-to-Italian localization; adapt evidence into Italian scripts, visualizations and short-form content while preserving terminology, meaning and source context.
-   Designed and built entropyforlife.it in WordPress; formally acknowledged in Giacomo Moro Mauretto's Mondadori book Italiani veri for scientific-literature research and error detection.

AI Evaluation & Model Behavior AnalystPage 1 of 2

## Additional relevant experience

### Founder & Project Lead

Yourself to Science™

[Website ↗](https://yourselftoscience.org)[Repository ↗](https://github.com/yourselftoscience/yourselftoscience.org)

Aug 2024 — Present

-   Founded and operate an open-source directory indexing more than 55 clinical studies, biobanks, donation programs, registries and other research initiatives.
-   Defined inclusion criteria, verification workflows, metadata structure, provenance requirements and licensing boundaries.
-   Define requirements, inspect code structure and behavior, test implementations and guide AI-assisted technical iteration through deployment and maintenance.

## Additional evidence

### [Research-integrity product requirements and testing](https://mariomarcolongo.com/notandia)

Created and operate Notandia (formerly MDPI Filter), a browser and Zotero research-integrity tool. It helps researchers identify articles from publishers whose editorial and peer-review practices have attracted scrutiny—including MDPI and Frontiers—and checks Crossref/Retraction Watch records for formal notices such as retractions, corrections and expressions of concern. I define the evidence rules, privacy safeguards, ambiguity handling, false-positive boundaries and release tests.

### [Structured evaluation practice](https://mariomarcolongo.com/security)

Repeated adversarial testing across instruction hierarchy, multimodal inputs, tool-use behavior and untrusted external context.

### [Investigation: evidence and interpretation](https://mariomarcolongo.com/integrity)

Attributed public cases demonstrate source synthesis, privacy-policy analysis, corporate-source reconciliation, archive recovery, legal-stage chronology, source-quality auditing, scientific taxonomy design and explicit separation of evidence from inference.

## Capabilities

**AI evaluation & adversarial testing**Exploratory adversarial testing, prompt and jailbreak analysis, multi-turn behavior, multimodal inputs, agentic tool-use and indirect prompt injection

**Evaluation planning & reporting**Test planning, evidence capture, reproducibility notes, taxonomy thinking, issue classification, severity-oriented reporting and mitigation-retesting concepts

**OSINT and research verification**Public-source and bibliographic research, claim decomposition, web-archive recovery, source-provenance analysis, corporate and legal record reconciliation, cross-source corroboration and reporting that separates evidence from inference

**Technical delivery**Codebase reading and behavior inspection, requirements definition, functional testing, AI-assisted implementation, Git/GitHub, JSON, REST APIs, WordPress, Cloudflare Pages and AWS Lambda deployment

**Multilingual quality and communication**Italian-native and English-C1 content review, English-to-Italian scientific localization, terminology consistency, source-faithful adaptation and clear evidence limitations for technical and non-technical audiences

## Best-fit role families

-   AI evaluation and adversarial testing
-   AI content red teaming and adversarial QA
-   Model behavior, trust & safety and policy testing
-   Human-data quality, grading and evaluation support

## Education & credentials

-   Information Technology — Open UAS Path Studies (120 ECTS) — Metropolia University of Applied Sciences, Aug 2026–Present; bachelor's-level ICT studies in software development, databases, Unix/Linux, data structures and algorithms, networking, cloud computing, cybersecurity and applied AI. Eligible to apply to the BEng IT programme after completing 120 ECTS (admission not yet granted).
-   GALENOS Crowd Evidence Synthesis Training — Cochrane Crowd & GALENOS, 2026
-   Career Essentials in Generative AI — Microsoft & LinkedIn, 2024
-   Italian — native. English — C1 overall (EF SET 68/100); advanced technical reading and professional writing, with practical English-to-Italian scientific localization and cross-language content-quality work.

Mario Marcolongo · me@mariomarcolongo.comPage 2 of 2
