Xuyang Cao

Xuyang Cao

Senior Algorithm Engineer, Ph.D.

Member, JD Tech Genius Team (TGT)

JD.com, Beijing, China
No. 18 Kechuang 11 Street, Tongzhou District, Beijing

Email: newxuyangcao [at] gmail [dot] com

[github] [google scholar]


Biography

Dr. Xuyang Cao is a senior algorithm engineer at JD Health Inc. in JD.com, specializing in multimodal large language models and their clinical applications. His current work focuses on agentic 3D medical image understanding and unified medical multimodal model training, with an emphasis on radiology-style reasoning, multimodal grounding, and deployable medical AI systems.

Since 2016, he has been engaged in research in computer vision, image processing, and pattern recognition. He completed his undergraduate, and Ph.D. degrees at Beijing Jiaotong University. During his Ph.D., his research focused on semantic segmentation and semi-supervised learning, under the supervision of professor Houjin Chen and professor Yanfen Li. Additionally, he closely collaborated with professor Yahui Peng during his doctoral studies.

Experience

Algorithm Engineer,  JD Health Inc.,   JD.com 2022-Present
Ph.D. Candidate,  School of Electronic and Information Engineering,   Beijing Jiaotong University 2017-2022
Master Candidate,  School of Electronic and Information Engineering,   Beijing Jiaotong University 2016-2017
Bachelor of Engineering,  School of Electronic and Information Engineering,   Beijing Jiaotong University 2012-2016

Projects

Agentic Strategy for 3D Medical Image Multimodal Understanding 2026.04-Present
  • Led the design of a ReAct-style agentic understanding system for medical visual question answering and radiology report generation, addressing the difficulty of directly processing 3D medical images with general-purpose MLLMs.
  • Simulated radiologists' reading workflow by building a coarse-to-fine reasoning mechanism covering continuous slice browsing, global screening, local close reading, and evidence integration.
  • Designed an agentic workflow that coordinates general MLLMs, medical expert models, and image-processing tools through multi-series reconnaissance, key-series selection, tool invocation, evidence aggregation, and structured output.
  • Enabled key lesion localization and cross-slice evidence acquisition for complex 3D medical image understanding scenarios.
  • On real-world medical benchmarks, the 27B agentic system outperformed GPT-5.5/5.6 bare-model baselines, improving report generation and medical VQA by +0.47pp and +6.15pp, respectively; related capabilities were deployed on the 2B MedWork platform and released at JD Discovery 2026.
Unified Medical Multimodal Foundation Model Training 2025.07-2026.04
  • Designed and trained an industry-leading medical omni-modal architecture across 7B/32B/72B scales, unifying medical image understanding, detection, segmentation, text, imaging, laboratory tests, and other clinical modalities.
  • Led model architecture design and partial model merging/training, focusing on semantic- and pixel-level multimodal understanding, Think with Grounding, and native 3D medical image understanding.
  • Demonstrated that unified multimodal models can surpass expert segmentation models for the first time, improving average performance by 112%.
  • Aligned chain-of-thought (CoT) reasoning with clinical diagnostic pathways, enabling localizable and verifiable multimodal reasoning.
  • Extended the framework to directly understand 3D medical images, proposed a visual token compression strategy, and introduced organ-level Think with Grounding, outperforming existing models on multiple benchmarks.
  • Achieved leading results across 7 task categories and 21 standard benchmarks, improving over SOTA models by +20.02pp in OCR, +8.71 in report generation, +7.44pp in 2D understanding, +8.09pp in 3D image understanding, +0.82pp in text QA, +23.70pp in detection, and +8.44pp in segmentation.
  • Accumulated million-scale medical multimodal data assets and released group-level open-source assets (Code, Technical Report); Jingyi Qianxun 2.0 was released at JD Discovery 2025 (link).
Digital Human Generation Based on Diffusion and Disentangled Facial Representation 2024.01-2025.05
  • Built JD Health's digital human generation technology from 0 to 1, including a GAN-based dual-tower model and a diffusion-based end-to-end generation model.
  • Proposed a real-time diffusion-based digital human generation method that disentangles dynamic facial expressions from static 3D identity information, trains a diffusion Transformer to generate identity-agnostic motion sequences from audio, and uses a dedicated renderer for high-quality animation synthesis.
  • Developed one of the industry's early models capable of audio-driven animation for both human portraits and animals, achieving an RTF of 0.76 and more than 40x faster inference than most existing diffusion models under the same hardware conditions.
  • Accumulated TB-scale video data assets, produced two types of digital human generation models and technical reports, and released open-source projects JoyVASA and JoyHallo, which together received 1,300+ GitHub stars.
  • Enabled online medical consultation, medical knowledge popularization, and representative products such as Dr. Dawei; related technologies were released at JD Discovery 2024.
Vision Technology for Medical Scenarios 2022.11-2024.01
  • Served as a core researcher at JD Health Exploration Research Institute, exploring classical computer vision technologies in medical imaging and online consultation scenarios, including early diagnosis of Alzheimer's Disease (AD) and Bipolar Disorder (BD) based on medical imaging.
  • Focused on explainable AI, multimodal fusion models, semantic segmentation, classification, and semi-supervised learning for medical applications.
  • Published 5 papers and multiple patents, with several brain imaging studies achieving leading results; skin disease diagnosis technology was deployed in online consultation scenarios.
  • Participated in a National Key R&D Program project with a 15-million RMB budget: Virtual Reality Cognitive Rehabilitation Training System for the Elderly.
Research on Deep Learning based Breast Mass (Ultrasound) and Lung Mass (MRI) Segmentation 2018.09-2022.04
  • Research on deep learning based semantic segmentation algorithms on breast ultrasound images as well as lung MRI images.
  • Delved deep into semantic segmentation algorithms, such as fully supervised learning, semi-supervised learning, neural network architecture search, etc.
  • Proposed lightweight dilated densely connected network for 3D breast tumor segmentation. The performance improved over 5% compared with classical segmentation networks, while network parameters were over 20 times smaller than classical networks.
  • Designed an uncertainty-aware temporal-ensembling model for semi-supervised segmentation. The semi-supervised method achieved 94.4% of the performance of supervised segmentation with only 1.1% labeled data.
  • Suggested an NAS-based 3D medical image segmentation framework, and achieved an improvement of 4.2% compared with the state-of-the-art human-designed segmentation network.
  • Related journal and conference papers have been published, total impact factor 30+.
High Speed Train Gear Defect Detection Based on Computer Vision 2018.09-2022.04
  • As a core developer, designed an automatic detection and quantitative analysis solution for surface pitting of coupling gear components in train sets, based on computer vision technology.
  • Collected coupling gear surface images and applied image recognition algorithms to generate statistics, reports, and alerts for gears exceeding the defect threshold.
  • Led the development of algorithms for detecting and segmenting gear surface wear areas, as well as the design and implementation of the software system on the Windows platform.
  • Completed defect detection of 104 gear surfaces (internal and external) with defects larger than 1mm within 3 minutes; published related papers and patents.
Geometric Parameters Measurement of an Overhead Line System 2018.09-2022.04
  • Provided a solution for measuring the geometric parameters of an overhead line system using scale factors and frame differences.
  • I was responsible for the preliminary algorithm simulation and participated in the hardware structure design work.
  • Two related patents have been granted.

Publications

G. Wang, Y. Li, Z. Zhou, S. An, X. Cao, Y. Jin, et al. PlgFormer: parallel extraction of local-global features for AD diagnosis on sMRI using a unified CNN-transformer architecture, Frontiers in Neurology 16 (2025): 1626922. [paper]

G. Wang, J. Zhao, X. Liu, Y. Liu, X. Cao, C. Li, Z. Liu, Q. Sun et al. Citrus-V: Advancing Medical Foundation Models with Unified Medical Image Grounding for Clinical Reasoning. Arxiv, 2025. [paper] [code][project]

J. Fei, SH. Ding, J. Zhao, J. Luo, G. Wang, Y. Yao, X. Cao, et al. Chronic disease trajectory network and Monte Carlo simulation in Chinese population. European Heart Journal, Volume 46, Issue Supplement_1, November 2025, ehaf784.4556. [paper]

J. Fei, G. Wang, J. Zhao, J. Luo, C. Zhang, Z. Wu, M. Gao, X. Cao, et al. Leveraging e-commerce user behavior data for common chronic diseases prediction. European Heart Journal, Volume 46, Issue Supplement_1, November 2025, ehaf784.4438. [paper]

G. Wang, X. Cao, S. An, F. Fan, C. Zhang, J. Wang, F. Yu, Z. Wang. Multi-Dimension-Embedding-Aware Modality Fusion Transformer for Psychiatric Disorder Classification, ICIGP, 2025. [paper]

X. Cao, G Wang, S Shi, J Zhao, Y Yao, J Fei, M Gao. JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation. Arxiv, 2024. [paper] [code][project]

S. Shi, X. Cao, J. Zhao, G. Wang. JoyHallo: Digital human model for Mandarin. Arxiv, 2024. [technical report] [code][project]

Z. Gao, Y. Guo, G. Wang, X. Chen, X. Cao, C. Zhang, S. An, F. Xu. Robust deep learning from incomplete annotation for accurate lung nodule detection. Computers in Biology and Medicine, 2024, 173:108361. [paper]

X. Cao, H. Chen, Y. Li, Y. Peng, Y. Zhou, L. Cheng, T. Liu, D. Shen. Auto-DenseUnet: Searchable Neural Network Architecture for Tumor Segmentation in 3D Automated Breast Ultrasound. Medical Image Analysis, 2022, 82: 102589. [paper]

Y. Zhou, H. Chen, Y. Li, X. Cao, S. Wang, D. Shen. Cross-Model Attention-Guided Tumor Segmentation for 3D Automated Breast Ultrasound (ABUS) Images. IEEE Journal of Biomedical and Health Informatics, 2022, 26(1): 301-311. [paper]

X. Cao, H. Chen, Y. Li, Y. Peng, S. Wang, L. Cheng. Uncertainty Aware Temporal-Ensembling Model for Semi-supervised ABUS Mass Segmentation. IEEE Transactions on Medical Imaging, 2021, 40(1):431-443. [paper]

X. Cao, H. Chen, Y. Li, Y. Peng, S. Wang, L. Cheng. Dilated Densely Connected U-Net with Uncertainty Focus Loss for 3D ABUS Mass Segmentation. Computer Methods and Programs in Biomedicine, 2021, 209: 106313. [paper]

J. Li, H. Chen, Y. Li, Y. Peng, N. Cai, X. Cao. AMRSegNet: Adaptive Modality Recalibration Network for Lung Tumor Segmentation on Multi-Modal MR Images. Multimedia Tools and Applications, 2021, 80: 33779–33797. [paper]

X. Cao, H. Chen, Y. Li, Y. Peng, Y. Zhou, L. Cheng. Boundary Loss with Non-Euclidean Distance Constraint for ABUS Mass Segmentation. 2020 CISP-BMEI, Chengdu, China, 2020, pp: 645-650. [paper]

Y. Peng, X. Cao, H. Chen, Y. Li, J. Li, X. Wang. Preliminary Study on Noise and Artifact Reduction in Phase-Contrast CT Image of Tristructural-Isotropic Coated Fuel Particle (in Chinese). Acta Electronica Sinica, 2019, 47(2): 448-453. [paper]

C. Wang, F. Li, Y. Li, H. Chen and X. Cao. A Defect Status Detecting Method for External Gear in Railway. 2018 IEEE 3rd International Conference on Image, Vision and Computing (ICIVC), Chongqing, 2018, pp: 123-127. [paper]

Y. Li, X. Cao, H. Chen, L. Zhang, N. Yang. Defect Status Detection Method Based on Machine Vision for External Gear in Train (in Chinese). Journal of The China Railway Society, 2018, 40(12):33-41. [paper]

J. Wei, X. Cao, H. chen, Y. li. Research on benign and malignant masses classification in mammogram (in Chinese). Journal of Beijing Jiaotong University, 2017, 41(5): 73-. [paper]

Patents & Books

Y. Peng, W. Jiang, Z. Zhu, H. Yang, X. Cao, H. Chen. A method of Measuring the Geometric Parameters of an Overhead Line System by using Geometric Magnification and Monocular Vision. China, CN201810182553.1, 2018-11-13. [Link]

Y. Peng, C. Zhang, B. Zheng, J. Yin, X. Cao, H. Chen. A method and a Device for Measuring the Geometric Parameters of an Overhead Line System by using Scale Factors and Frame Differences. China, CN201710464403.5, 2017-06-19. [Link]
Y. Zhou, X. Cao. Neural Networks with TensorFlow 2, Apress, 2020.   [translated] [Link]

Personal Qualifications

Research Interests: Multimodal Large Language Models, Agentic 3D Medical Image Understanding, Unified Medical Multimodal Modeling
Training & Deployment: SFT, GRPO, DeepSpeed, FlashAttention, Distributed Training, Visual Token Compression, Model Evaluation and Serving
Hobbies: Personal blog with over 330,000 PV  |  Enjoy reading, hiking, and traveling