About

Hi, I'm Yiwen Sun 👋

PhD candidate in Computer Science (AI) at The Hong Kong Polytechnic University, researching medical image segmentation under Prof. Cai Jing.

Education

PhD · The Hong Kong Polytechnic University · Computer Science (AI, Image Segmentation) · Supervisor: Prof. Cai Jing2024.01 – 2027.01
MSc · University of Birmingham · Electronic and Computer Engineering2022.09 – 2023.09
BEng · Beijing Institute of Technology · Measurement and Control Technology (Optics)2017.09 – 2021.06

Featured Projects

🗂 Multimodal Dataset Factory · mm_dataset_factory Details →
One-click pipeline from Word exam papers to multimodal error-question image datasets. 4-stage pipeline (structuring → VLM error planning → AI image gen → quality gate) + SSE streaming + post-generation evaluation loop. Ships with a full bilingual docs repo.
📊 The datasets it produces became a benchmark → EduFig-IC: graded image–text consistency
FastAPIReact 19SQLAlchemy asyncSSEMultimodal
⚙️ Multilingual Exam Review Agent · exam-checker-agent Details →
Auto question splitting + per-question routed detection + four-stage image-text review (text-only re-review for recall) + three-way plagiarism + measurable evaluation; React + FastAPI + Docker full-stack.
ReactFastAPISQLiteDockerFull-stack
More projects →

Research Experience

Exam Review & Multimodal Data Engineering · Internship (Talkweb)2026.03 – 2026.06
LLM / Agent / Multimodal / Data Engineering
  • Exam Review Agent Platform · exam-checker-agent: built a platform that reviews university final and postgraduate exam papers, auto-detecting 10 error classes — typos, grammar, logic, semantics, image-text consistency, etc. (20 detection tasks across ZH/JA/DE/FR); a "parse → per-question routing → detection → dedup & aggregation → plagiarism check → report" pipeline, with a four-stage image-text chain (VLM visual check + forced reconciliation checklist + text-only re-review) that raised image-text inconsistency recall from 52% to 88%, F1 0.89–0.92, precision 0.95+. Full-stack Python / FastAPI / React / SQLite / Docker.
  • Review Test-Dataset Generator · mm_dataset_factory: built a platform that synthesizes test data for the review system — extracting figure-bearing questions from real Word/DOCX papers, planning errors at three levels L1/L2/L3 (image mismatch / parameter conflict / logically unsolvable) via VLM error planning + Seedream / Wan text-to-image models for error injection + a post-generation evaluation loop (numeric-anchor check, pass/repair/regenerate), batch-producing image-text pairs where the figure contradicts the stem. Full-stack FastAPI + React + SQLAlchemy async, SSE streaming.
  • Image-Text Consistency Benchmark · EduFig-IC: data produced by the pipeline above, curated into a graded image-text consistency benchmark for STEM questions with figures (L1/L2/L3, 973 samples), evaluating VLMs' ability to detect image-text inconsistency.
Multimodal Super-Resolution Agent · Paper Reproduction & Extension2025.12 – 2026.01
Python / PyTorch / vLLM
  • Built an SR Agent on open-source frameworks and vLLM, with a full offline inference chain: planning → tool calls → image restoration.
  • Rewrote planning prompts for the local Llama-Vision model, reducing over-reasoning and improving executable plans and tool-call stability.
Zero-Annotation Pathology Nuclei Segmentation · First Author2024.01 – 2024.12
Pattern Recognition Letters (SCI / JCR Q2) · Python / PyTorch
  • A pretrained Vision-Language detector produces zero-shot coarse boxes, enabling training without any pathology annotation.
  • Designed Gaussian-confidence weak supervision + focal loss + Voronoi/k-means complementary coarse labels driving fine nucleus-boundary learning.
Topology-Aware Segmentation of Tubular Structures in 3D Microscopy · First Author2025.01 – 2025.12
Physics in Medicine & Biology (SCI / JCR Q1, Under Review) · Python / PyTorch / 3D Seg
  • Designed a radius-field topology loss + a large-kernel 3D U-Net, improving connectivity and Dice by 4%.
  • Defined annotation and quality-control workflows with partners; fine-tuned and evaluated on a private dataset.

Publications & Talks

  1. Yiwen Sun, Ranran Zhang, Fuqiang Chen, Kun Ru, Miaoxia He, Qizhai Li, Yao Pu, Jing Cai, Wenjian Qin. “ZA-Net: A universal zero-annotation nuclei segmentation network for pathology images via vision-language pre-trained model.” Pattern Recognition Letters, 2026 (SCI / JCR Q2) · First author · DOI
  2. Fuqiang Chen, Ranran Zhang, Wanming Hu, Deboch Eyob Abera, Yue Peng, Boyun Zheng, Yiwen Sun, Jing Cai, Wenjian Qin. “PGVMS: A Prompt-Guided Unified Framework for Virtual Multiplex IHC Staining with Pathological Semantic Learning.” IEEE Transactions on Medical Imaging (SCI / JCR Q1) · Co-author
  3. Fuqiang Chen, Ranran Zhang, Boyun Zheng, Yiwen Sun, Jiahui He, Wenjian Qin. “Pathological Semantics-Preserving Learning for H&E-to-IHC Virtual Staining.” MICCAI 2024 (CCF B) · Co-author
  4. Invited talk at the 3rd Intelligent Medicine Symposium

Skills / Strengths & Others

  • Programming: Python, PyTorch, Linux; LLM/VLM inference deployment (vLLM), Agent systems and data-pipeline engineering.
  • Huawei HarmonyOS app: published a health app (food recognition / calorie tracking), 2,000+ downloads; 400+ class MobileNet offline inference.
  • Language: English, IELTS 6.5.
  • Awards: 2023 UoB PGR Research Image Runner-up; 2019 BIT Outstanding Student & Academic Progress Scholarships.
  • Interests: photography (VCG contracted photographer); hiking; content creation (1,000+ followers on Xiaohongshu & Bilibili).