Healthcare — New Datasets 23
7 authorized medical journal articles, covering illustrated medical literature, clinical case reports, and medical imaging captions with full reprint authorization.
500 real structured health checkup datasets with balanced gender distribution (1:1 male-to-female), multiple age segments, and covering individuals with checkup amounts above 1,000 CNY.
The largest medical NLP training corpus, featuring 5,000 de-identified electronic medical records covering chief complaints, history of present illness, past medical history, admission records, discharge summaries, and more.
Covering a massive collection of authoritative medical knowledge content including medical textbooks, clinical guidelines, review articles, and drug formularies.
500 sets of 12-lead ECG waveform images and diagnosis reports, including automated analysis reports and physician-reviewed diagnoses.
Covering dermoscopy images, clinical photos of skin lesions, and pathological gold-standard diagnosis reports.
500 chest X-ray DICOM images with structured diagnosis reports, dual-reviewed to gold-standard quality.
500 head CT plain and enhanced DICOM series with diagnosis reports, covering cerebral infarction, cerebral hemorrhage, intracranial tumors, and more
High-quality fundus color photographs and structured diagnostic reports, with ICDR international standards for diabetic retinopathy staging
Whole-slide digital images (20x+ scan magnification) with complete pathological diagnosis reports, including immunohistochemistry results and TNM staging.
500 endoscopy videos with key frame captures and diagnosis reports, including polyp location, size, and type
500 upper abdominal multi-organ cancer annotations, covering the liver, gallbladder, pancreas, spleen, adrenal glands, kidneys, and intraperitoneal lymph nodes.
500 full diagnostic process pathology image+text datasets, covering the complete diagnostic chain
500 difficult case MDT consultation texts for rare and complex disease AI training - scarce, high-difficulty case data.
100 longitudinal chronic disease tracking structured datasets, with 7+ consecutive inpatient visits tracked.
500 Qupath fine-grained region annotations with 6 label types: HCC, ICC, MID, TLS, and MVI.
100 liver lesion image annotations for differentiating liver cancer from hemangioma.
100 breast tumor image annotations for breast cancer AI screening and diagnosis.
200,000+ full-cycle inpatient record texts and PDFs, covering medical record quality control, DRG, and DIP
5,000,000+ medical entities and 20,000,000+ relationship triplets, with ICD-10/11 and SNOMED CT mapping across 7 major knowledge bases
11,000,000+ entries of professional medical NLP annotated corpus, covering 12 task types including entity extraction, conversational understanding, and chain-of-thought, across 66 sub-datasets
2,000,000+ medical imaging training samples, fully covering CT, MRI, X-ray, ultrasound, endoscopy, and pathology, with classification, detection, and segmentation annotations
3,000,000+ specialized disease structured datasets, covering 26 key disease types including diabetes, stroke, and cancer, with 5+ years of long-term follow-up
Healthcare — Existing Datasets 11
1,000,000+ in-depth health checkup reports with 800+ structured medical fields
328,000 entries of electronic medical records, diagnosis reports, medical orders, and literature summaries for NER and relation extraction annotation
52,800 multi-sequence HCC case MRI DICOM images with professional physician annotations
General medical imaging annotation dataset covering four major modalities: CT, X-Ray, Ultrasound, and MRI
1,000,000+ annotated surgical video frames with 24-type pixel-level semantic segmentation of anatomical structures and surgical instruments
Hundreds of thousands of obstetrics and gynecology clinical datasets, covering prenatal emergencies, high-risk pregnancies, and full pregnancy scenarios including delivery and postpartum
1,000,000+ tongue surface and facial images with 50,000+ professional annotations
1,000,000+ DICOM image slices covering 7 major organs, thin-slice CT at 512x512+ resolution
Pharma NEW 6
Newly launched, covering pharmaceutical structured data, TCM formulas, medical devices, and consumables pharma data assets.
5,000,000+ entries of China pharmaceutical data with 15+ core dimensions, full ATC Level 5 coverage, supporting rational drug use AI and pharmacovigilance
2,000,000+ entries of TCM, formulas, and proprietary medicine data, covering nature and flavor, meridian tropism, efficacy classification, formula composition, and genuine production regions
1,000,000+ entries of medical device registration and consumables data, covering registration numbers, management categories, scope of application, and technical specifications
Insurance NEW 4
Newly launched, covering insurance settlement lists, DIP disease type grouping, DRG grouping codes, and insurance payment standard data.
1,000,000+ entries of insurance settlement lists and payment data with 14+ core dimensions, covering insurance catalogs, DIP grouping, DRG coding, and cost structure
Intl Corpus NEW 1
Newly launched, covering 500,000+ multilingual medical corpus entries across 10 languages, supporting medical neural machine translation and cross-lingual knowledge graph construction.
500,000+ multilingual medical corpus entries across 10 languages and 8 major types, with ICD-11/SNOMED-CT terminology annotation at S/A-level translation quality