Open Ended Medical Reinforcement Learning
-
Updated
Mar 15, 2026 - Python
Open Ended Medical Reinforcement Learning
Synthetic medical VQA pipeline: 119K images annotated by frontier VLMs, cross-validated at 93% agreement, fine-tuned on 3 model families (2-3B params)
[ICMR'21, Best Poster Paper Award] Medical Visual Question Answering with Multi-task Pre-training and Cross-modal Self-attention
Official repository for the ACL 2025 Findings paper "Worse than Random? An Embarrassingly Simple Probing Evaluation of Large Multimodal Models in Medical VQA"
[IEEE TMI'22] VQAMix: Conditional Triplet Mixup for Medical Visual Question Answering
Visual Question Answering in the Medical Domain
Research framework evaluating retrieval-augmented vision-language models for evidence-grounded medical visual question answering.
AstraQ-VL is a minimal, efficient astronomy Vision-Language Model following the LLaVA architecture. Stage 1 trains only a lightweight MLP connector to align frozen CLIP vision features with a frozen LLM. Stage 2 then warm-starts that connector and fine-tunes the LLM with LoRA adapters on instruction (QA) data.
Giving eyes to a blind medical LLM: applying VoRA (Vision as LoRA) to BioMistral-7B for radiology Visual Question Answering on chest X-rays.
Healthcare Multimodal RAG system for medical visual question answering and retrieval-augmented reasoning using LLaVA, QLoRA, OpenI chest X-rays, and scalable modular architecture with future ColQwen2, BM25, CLIP, RRF, and Qwen2-VL integration.
MERGETUNE-inspired continued fine-tuning for cross-domain medical VQA transfer between VQA-RAD and PathVQA.
Vision-language fusion for clinical image QA. Cross-attention BioViL-T + Mistral-7B QLoRA, MC Dropout confidence, Grad-CAM explainability, dual local/API inference. GPT-4o baseline: 56% on VQA-RAD with 41% abstention.
Vietnamese Medical Visual Question Answering
Official repo for hedge-bench PyPI package for hallucination detection in vision-language VQA models, including the codebase for the paper “HEDGE: Hallucination Estimation via Dense Geometric Entropy for VQA with Vision-Language Models.”
A curated literature resource hub for Medical Visual Question Answering, covering surveys, datasets, benchmarks, evaluation metrics, representative methods, and multimodal medical agents, with a focus on the shift from passive answer prediction to active, evidence-seeking clinical inquiry.
A curated collection of gastrointestinal endoscopy datasets for AI research, covering colonoscopy, polyp detection, segmentation, classification, capsule endoscopy, surgical endoscopy, medical VQA, and multimodal learning.
Medical PRM pipeline for VQA datasets with commercial, open-model, and demo backends
Add a description, image, and links to the medical-vqa topic page so that developers can more easily learn about it.
To associate your repository with the medical-vqa topic, visit your repo's landing page and select "manage topics."