
Powered by Dr7.ai Medical AI Platform
BiomedCLIP is a CLIP-style model from Microsoft Research that maps medical images and text into one shared embedding space. It was pretrained on PMC-15M — 15 million figure-caption pairs from 4.4 million PubMed Central articles — and works zero-shot on classification, retrieval and visual question answering.
Upload a medical image and describe what you want to check — the demo pairs visual understanding with biomedical text reasoning in the same way BiomedCLIP aligns images and captions.
Experience Advanced Medical AI
Get unlimited image-text analysis, higher upload limits and API access to run BiomedCLIP-style retrieval and classification at scale.
BiomedCLIP is a biomedical vision-language foundation model released by Microsoft Research in 2023. It follows the CLIP recipe — contrastive pretraining of an image encoder and a text encoder — but replaces web data with scientific figures and their captions, so the model learns the visual vocabulary of radiology, pathology, microscopy and clinical photography.
The image encoder is a ViT-B/16 and the text encoder is PubMedBERT, the domain-specific BERT that already understands biomedical terminology. Training data is PMC-15M, a corpus of 15 million image-caption pairs mined from 4.4 million open-access PubMed Central articles — roughly two orders of magnitude larger than earlier biomedical multimodal datasets.
Because the two encoders share an embedding space, BiomedCLIP can classify images by comparing them with text prompts it has never been trained on ("zero-shot"), retrieve images from text queries, and serve as the visual backbone for medical VQA systems. Weights are openly available on Hugging Face under the MIT licence.
Released by Microsoft Research in 2023 (Zhang et al., "BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs")
Open weights (MIT licence) on Hugging Face; widely used as the image backbone for medical VQA and retrieval research
Contrastive image-text pretraining tuned for biomedical figures
Zero-shot medical image classification from text prompts
Image-to-text and text-to-image retrieval across modalities
Shared embeddings for radiology, pathology, microscopy and clinical photos
Cross-modal similarity scoring between an image and a caption
Drop-in visual encoder for medical visual question answering
Pretrained on 15M figure-caption pairs from PubMed Central
ViT-B/16 image encoder + PubMedBERT text encoder
Open weights under the MIT licence
Reported in the BiomedCLIP paper (Microsoft Research, 2023); numbers below are the tasks, not Dr7.ai measurements
Outperformed general CLIP and prior biomedical CLIP models on pathology, radiology and dermatology image sets without task-specific training
Best image-to-text and text-to-image retrieval on PMC-derived benchmarks at release
Used as the image encoder for state-of-the-art medical visual question answering pipelines
Pretraining corpus of 15 million image-caption pairs from 4.4 million PubMed Central articles
Retrieve medical images with free-text queries thanks to the shared image-text embedding space.
Label images by comparing them with text prompts — no fine-tuning required.
Build datasets by retrieving figures from the literature that match a textual description.
Os modelos descritos nesta página são ferramentas de pesquisa para imagem clínica. Se o que você tem é uma foto do dia a dia — uma erupção, uma pinta, uma unha, um ferimento —, o AI Doctor Scan a analisa e devolve gratuitamente um nível de urgência e um resumo em linguagem simples em cerca de 30 segundos. Não nomeia nenhum diagnóstico e não exige conta.
Experimente o AI Doctor Scan — grátis