AI evaluation · scientific fact-checking

I test AI systems and verify scientific claims.

My primary work is AI model-behavior evaluation and scientific fact-checking. I also build and operate research-information tools when a recurring verification problem needs a practical system.

Italy · EU work authorization · worldwide relocation · international contracting

Professional scope

Where I can contribute.

01 · Primary

AI model-behavior evaluation

Adversarial and edge-case testing across chat, multimodal inputs, tool use and indirect prompt injection.

02 · Primary

Scientific fact-checking

Primary-literature review for scripts, articles and public scientific communication.

03 · Supporting

Research information systems

Structured records, provenance rules and maintained tools for recurring research-information problems.

Selected work

Selected work, shown through the actual output.

Each case shows the actual public output, my contribution and the boundary of what that work demonstrates.

Screenshot of the dated Gray Swan Arena profile showing Proving Ground rank, percentile, breaks and separate Arena activity metrics#75Proving GroundTop 6%Global percentile110Total breaks#371Arena rank
01AI evaluation

Finding model failures across four evaluation surfaces.

Across four public evaluation surfaces, I test chat, image, agent and indirect prompt-injection behavior and preserve dated evidence of the platform-reported result.

What I did
I vary the interaction path, preserve reproduction notes and distinguish direct observations from platform labels, independent verification and model-wide conclusions.
Result
#75 on the Proving Ground, top 6%, with 110 platform-displayed total breaks on 26 July 2026; the Arena profile displayed rank #371, 246 submissions, 27 global unique breaks and 1,090 points.

Scope: The four visible area counters sum to 109 while the profile displays 110 total breaks. Both are reported without inferring the platform’s internal aggregation. This supports evaluation and adversarial-QA applications, not penetration-testing or senior red-team engineering claims.

Entropy for Life · paid contractor

Research, content and website work

I contributed to 80 documented published pieces.

80published contributions55 YouTube videos · 4 articles · 21 short-form pieces
Research and fact-checkingPrimary literature, source verification and evidence synthesis
Content productionScripts, data analysis, visualizations, slides and selected thumbnails
Website design and managementWordPress, responsive design, publishing and OVHcloud operations
YouTube channel
36.5Mviews
267K subscribers592 videos
Instagram 159KTikTok 54K · 528K likes
Official evidence indexVideos · articles · short-form work · selected thumbnails
View the official Entropy for Life work record
02Scientific content quality & operations

Evidence quality and content operations at creator scale.

Paid contractor supporting an established Italian science-communication brand across evidence review, content production and website operations.

What I owned
Recurring primary-literature research and scientific fact-checking. Depending on the assignment, I also developed scripts, data analyses, visualizations, slides, on-screen assets, short-form content and selected thumbnails. I designed and built entropyforlife.it in WordPress and manage its responsive design, publishing and OVHcloud technical operations.
Result
80 documented published content contributions: 55 YouTube videos, 4 co-authored articles and 21 short-form pieces. The official work record also indexes selected thumbnail work, which overlaps with video projects and is not added to the total.

Scope: Platform metrics describe the production environment, not a personal audience. Quantified thumbnail lift is stated only when comparable analytics are available.

Screenshot of the Yourself to Science research-participation directoryLive public product · catalogue, provenance and machine-readable records
03Supporting research system

Building a maintained directory of research opportunities.

Yourself to Science turns scattered institutional opportunities into a public catalogue with explicit inclusion and update rules.

What I did
I defined the inclusion criteria, verification fields, provenance model, licensing boundaries, update workflow and public-data requirements, then coordinated AI-assisted implementation and ongoing operation.
Result
More than 55 opportunities indexed, with FAIRsharing and Zenodo records and human- and machine-readable interfaces.

Scope: My contribution covers requirements, information architecture, verification, functional testing, deployment diagnosis and operations—not unaided software development.

Current product

MDPI Filter now works in the browser and as a Zotero plugin.

The current product identifies MDPI references across literature-search and reference-management workflows while avoiding ambiguous title-based matches. The broader rebrand and expansion to retractions, comments and other research-integrity signals are future work, not shipped functionality.

Chrome · Edge · Firefox · Safari source · Zotero 7–9
Scientific visualization

A diagram that became a reusable public reference.

I designed this vector diagram to show overlap among monogenic conditions associated with autism, dystonia, epilepsy and schizophrenia. The Wikimedia record exposes the source file, authorship, revision history and reuse across four Wikipedia language editions.

Open the Wikimedia source record
Euler diagram of overlapping clinical phenotypes in genes associated with autism spectrum disorder, dystonia, epilepsy and schizophrenia
More public work

Additional tools and experiments.

Working preferences

Clear goals, direct feedback and responsibility for the result.

I work best when the objective and decision rights are explicit, feedback is direct and I can follow a problem through investigation, documentation, release and maintenance.

01

Separate observation from interpretation

State what was directly observed before drawing a broader conclusion.

02

Make assumptions visible

Document the definitions, exclusions and judgments that affect the result.

03

Plan for maintenance

Treat updates, provenance and operational recovery as part of the work.

Experience

Relevant experience.

The targeted CVs carry the full application detail.

Model-Behavior Evaluator

Independent practice · Gray Swan Proving Ground participant · Conduct self-directed adversarial testing across chat, multimodal, agentic tool-use and indirect prompt-injection settings.

Scientific Fact-Checking, Content Operations & Audience Optimization Contractor

Entropy for Life — Italy · Delivered 80 documented published content contributions—55 YouTube videos, 4 co-authored articles and 21 short-form pieces—for an Italian science-communication brand with 267K YouTube subscribers and 480K+ combined platform following; recurring evidence review, content production, selected thumbnails and website operations.

Founder & Research-Workflow Owner

Yourself to Science™ · Founded and operate an open-source research-participation directory indexing more than 55 initiatives, with documented verification, provenance and metadata workflows.

Application documents

Start with the role you are hiring for.

The AI evaluation CV is the recommended default. The others are deliberately tailored alternatives.

Current focus

AI evaluation and scientific evidence roles.

Especially model-behavior evaluation, scientific fact-checking, research-data quality and knowledge-integrity work.