JMZhiao (Jacky) MoAI / Software / Robotics
HomeProjectsExperienceAboutResumeContact

Zhiao (Jacky) Mo

AI / Machine Learning · Software Engineering · Robotics & Computer Vision

Melbourne, Australia

zhiaomo@gmail.com
ProjectsExperienceAboutResumeContactLinkedInGitHub

Built with Next.js, three.js, Framer Motion and hand-made SVG.

Back to all projects
NLP · Information Retrieval · Research

Climate Claim Fact-Checking Pipeline

Hybrid retrieval + re-ranking + stance classification over a 1.2M-passage corpus

CC

Role

Team of 3 (COMP90042, University of Melbourne). The submitted system used my end-to-end pipeline; the conflict-aware evidence-selection strategy is an original contribution of the work.

When

Semester 1, 2026 · University of Melbourne

Tech

  • Python
  • PyTorch
  • Transformers
  • TF-IDF
  • all-MiniLM-L6-v2
  • ClimateBERT
  • RRF

A three-stage system that verifies climate-related claims against a very large evidence corpus: hybrid sparse + dense retrieval, transformer re-ranking, then pair-level stance verification — finished with a conflict-aware evidence-selection strategy built specifically to surface the DISPUTED class instead of just taking the top-ranked passages.

Numbers that are real

1,208,827 passages

Evidence corpus

0.908

Stage-1 any-hit @ top-500

0.279 → 0.506

Claim classification accuracy

0.056 → 0.444

DISPUTED recall

What I did

  • ✓Fused TF-IDF sparse retrieval with all-MiniLM-L6-v2 dense retrieval using Reciprocal Rank Fusion, reaching a 0.908 any-hit retrieval rate at a top-500 candidate pool.
  • ✓Fine-tuned climatebert/distilroberta-base-climate-f as a binary selector to re-rank the pool down to the top 150 passages.
  • ✓Fine-tuned a separate ClimateBERT model for pair-level stance verification, then built a conflict-aware greedy evidence selection that jointly scores relevance, stance diversity and redundancy to pick the final 4 passages.
  • ✓Ran systematic sweeps instead of defaulting to the final checkpoint: retrieval pool sizes from 50 to 5000, selector/verifier epochs, and final evidence-set size, which showed the trade-off between overall harmonic mean and minority-class recall.

Honest limitations

  • Evidence F-score settles at 0.167 — retrieval coverage is strong, but choosing the right evidence is still the bottleneck.
  • Overall accuracy (0.506) and minority-class DISPUTED recall still trade off against each other; the harmonic mean stays modest.
  • Evaluated offline against the course benchmark only — the pipeline was not deployed as a service.

Projects

Stock Forum Summarization & Sentiment Analysis System

Read case studyBack to all projects
NLP