🗂 Multimodal Dataset Factory · mm_dataset_factory
One-click pipeline from Word exam papers to multimodal error-question image datasets. 4-stage pipeline + SSE streaming + post-generation evaluation loop, with a full bilingual docs repo.
View details →
📊 Image–Text Consistency Benchmark · EduFig-IC
A graded image–text consistency benchmark over STEM figure-based exam questions (L1/L2/L3, 973 samples) for evaluating image–text alignment and hallucination detection in multimodal LLMs. Data from the dataset factory.
View benchmark →
⚙️ Multilingual Exam Review Agent · exam-checker-agent
Auto question splitting + per-question routed detection + four-stage image-text review (text-only re-review for recall) + three-way plagiarism + measurable evaluation; React + FastAPI + Docker full-stack.
View details →
📝 AI Exam Review System · exam-smart-checker
LLM + vector-retrieval platform for exam error detection and duplicate checking, with multi-dimension judgment, question-type routing and historical-bank comparison.
View details →
🧰 Review Toolkit · checker-tools
A 4-in-1 Streamlit toolbox covering PDF parsing, bad-case generation, model evaluation and training-set export, with Docker Compose deployment.
View details →