Plan Pro: 30% de descuentoObtener
Medical AI Background

BiomedCLIP

Biomedical Vision-Language Model (Microsoft)

Powered by Dr7.ai Medical AI Platform

BiomedCLIP is a CLIP-style model from Microsoft Research that maps medical images and text into one shared embedding space. It was pretrained on PMC-15M — 15 million figure-caption pairs from 4.4 million PubMed Central articles — and works zero-shot on classification, retrieval and visual question answering.

🔗
CLIP
Medical AI Model
15M
Pairs
Microsoft Research
Vision-Language Model

Try BiomedCLIP-Style Image-Text Analysis

Upload a medical image and describe what you want to check — the demo pairs visual understanding with biomedical text reasoning in the same way BiomedCLIP aligns images and captions.

BiomedCLIP Demo

Experience Advanced Medical AI

0/3 messages

Start Medical Consultation

Ask complex medical questions or seek clinical reasoning assistance

Desbloquea todo el potencial de BiomedCLIP

Get unlimited image-text analysis, higher upload limits and API access to run BiomedCLIP-style retrieval and classification at scale.

CLIP
CLIP Architecture
15M
Image-Text Pairs
10+
AI Models Available
Actualizar a Pro→

What is BiomedCLIP?

BiomedCLIP is a biomedical vision-language foundation model released by Microsoft Research in 2023. It follows the CLIP recipe — contrastive pretraining of an image encoder and a text encoder — but replaces web data with scientific figures and their captions, so the model learns the visual vocabulary of radiology, pathology, microscopy and clinical photography.

The image encoder is a ViT-B/16 and the text encoder is PubMedBERT, the domain-specific BERT that already understands biomedical terminology. Training data is PMC-15M, a corpus of 15 million image-caption pairs mined from 4.4 million open-access PubMed Central articles — roughly two orders of magnitude larger than earlier biomedical multimodal datasets.

Because the two encoders share an embedding space, BiomedCLIP can classify images by comparing them with text prompts it has never been trained on ("zero-shot"), retrieve images from text queries, and serve as the visual backbone for medical VQA systems. Weights are openly available on Hugging Face under the MIT licence.

🔗

Latest Development

Released by Microsoft Research in 2023 (Zhang et al., "BiomedCLIP: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs")

Zero
Shot

Open weights (MIT licence) on Hugging Face; widely used as the image backbone for medical VQA and retrieval research

Key Features

Contrastive image-text pretraining tuned for biomedical figures

Core Capabilities

Zero-shot medical image classification from text prompts

Image-to-text and text-to-image retrieval across modalities

Shared embeddings for radiology, pathology, microscopy and clinical photos

Cross-modal similarity scoring between an image and a caption

Drop-in visual encoder for medical visual question answering

Pretrained on 15M figure-caption pairs from PubMed Central

ViT-B/16 image encoder + PubMedBERT text encoder

Open weights under the MIT licence

Performance Benchmarks

Reported in the BiomedCLIP paper (Microsoft Research, 2023); numbers below are the tasks, not Dr7.ai measurements

Zero-shot classification

SOTA

Outperformed general CLIP and prior biomedical CLIP models on pathology, radiology and dermatology image sets without task-specific training

Cross-modal retrieval

SOTA

Best image-to-text and text-to-image retrieval on PMC-derived benchmarks at release

VQA-RAD / SLAKE

Backbone

Used as the image encoder for state-of-the-art medical visual question answering pipelines

PMC-15M

15M pairs

Pretraining corpus of 15 million image-caption pairs from 4.4 million PubMed Central articles

Casos de Uso

🔍

Medical Image Search

Retrieve medical images with free-text queries thanks to the shared image-text embedding space.

🏷️

Zero-shot Classification

Label images by comparing them with text prompts — no fine-tuning required.

🔬

Biomedical Research

Build datasets by retrieving figures from the literature that match a textual description.

¿Tienes una foto de un síntoma visible?

Los modelos descritos en esta página son herramientas de investigación para imagen clínica. Si lo que tienes es una foto cotidiana —una erupción, un lunar, una uña, una herida—, AI Doctor Scan la analiza y devuelve gratis un nivel de urgencia y un resumen en lenguaje sencillo en unos 30 segundos. No nombra ningún diagnóstico y no requiere cuenta.

Prueba AI Doctor Scan — gratis