Jiahao Zhang

Research Scientist @ SB Intuitions · Ph.D., The University of Osaka

I am a Research Scientist at SB Intuitions, working on research and development for Responsible AI. I received my Ph.D. from the D3 Center (formerly the Institute for Datability Science), The University of Osaka, where I was advised by Prof. Yuta Nakashima and Prof. Hajime Nagahara, and my M.S. from the Graduate School of Medicine, Osaka University.

My research interests include Responsible AI, visual in-context learning, and the safety and alignment of vision–language models. I am also interested in medical image analysis and AI applications for healthcare.

Portrait of Jiahao Zhang

News

Selected Recent Work

A closer look at three recent projects. * denotes first author.

PANICL overview
arXiv

Visual In-Context Learning Training-free

PANICL: Mitigating Over-Reliance on Single Prompt in Visual In-Context Learning

Jiahao Zhang*, Bowen Wang, Hong Liu, Yuta Nakashima, Hajime Nagahara

Visual in-context learning conditions on a single input–output image pair, which makes predictions biased and unstable. PANICL is a training-free framework that aggregates multiple in-context pairs instead, smoothing assignment scores across them. It improves foreground segmentation, single-object detection, colorization, multi-object segmentation and keypoint detection over strong baselines, and holds up under both dataset-level and label-space domain shifts.

InMeMo overview
WACV 2024

Visual In-Context Learning Prompt Learning

Instruct Me More! Random Prompting for Visual In-Context Learning

Jiahao Zhang*, Bowen Wang, Liangzhi Li, Yuta Nakashima, Hajime Nagahara

Visual in-context learning hands a large frozen model a single input–output image pair to exemplify the task, so accuracy hinges on how good that pair is. InMeMo augments the in-context pair itself with a learnable perturbation, letting the demonstration adapt while the model stays frozen. This lightweight training lifts mIoU by 7.35 on foreground segmentation and 15.13 on single-object detection over the no-learnable-prompt baseline.

DiReCT pipeline
NeurIPS 2024

LLM Reasoning Benchmark

DiReCT: Diagnostic Reasoning for Clinical Notes via Large Language Models

Bowen Wang, Jiuyang Chang, Yiming Qian, Guoxin Chen, Junhao Chen, Zhouqiang Jiang, Jiahao Zhang, Yuta Nakashima, Hajime Nagahara

LLMs answer medical exam questions well, but rarely show reasoning a clinician can audit. DiReCT is a benchmark of 511 clinical notes, each annotated by physicians with the full diagnostic path from observation to final diagnosis, paired with a diagnostic knowledge graph. Evaluating leading LLMs on it exposes a wide gap between their reasoning and that of human doctors.

Publications

Selected works; see Google Scholar for the full list. * denotes first author.

International Conferences

DiReCT overview
NeurIPS 2024

DiReCT: Diagnostic Reasoning for Clinical Notes via Large Language Models

Bowen Wang, Jiuyang Chang, Yiming Qian, Guoxin Chen, Junhao Chen, Zhouqiang Jiang, Jiahao Zhang, Yuta Nakashima, Hajime Nagahara

InMeMo overview
WACV 2024

Instruct Me More! Random Prompting for Visual In-Context Learning

Jiahao Zhang*, Bowen Wang, Liangzhi Li, Yuta Nakashima, Hajime Nagahara

Journal Articles

ILD-Slider
J. Imaging

ILD-Slider: A Parameter-Efficient Model for Identifying Progressive Fibrosing Interstitial Lung Disease from Chest CT Slices

Jiahao Zhang*, Shoya Wada, Kento Sugimoto, Takayuki Niitsu, Kiyoharu Fukushima, Hiroshi Kida, Bowen Wang, Shozo Konishi, Katsuki Okada, Yuta Nakashima, Toshihiro Takeda

E-InMeMo overview
J. Imaging

E-InMeMo: Enhanced Prompting for Visual In-Context Learning

Jiahao Zhang*, Bowen Wang, Hong Liu, Liangzhi Li, Yuta Nakashima, Hajime Nagahara

CXR annotations
CMPB

Automatic Creation of Annotations for Chest Radiographs Based on the Positional Information Extracted from Radiographic Image Reports

Bowen Wang, Toshihiro Takeda, Kento Sugimoto, Jiahao Zhang, et al.

Preprints

PANICL overview
arXiv

PANICL: Mitigating Over-Reliance on Single Prompt in Visual In-Context Learning

Jiahao Zhang*, Bowen Wang, Hong Liu, Yuta Nakashima, Hajime Nagahara

VASCAR overview
arXiv

VASCAR: Content-Aware Layout Generation via Visual-Aware Self-Correction

Jiahao Zhang*, Ryota Yoshihashi, Shunsuke Kitada, Atsuki Osanai, Yuta Nakashima

E-SCOUTER overview
arXiv

Explainable Image Recognition via Enhanced Slot-Attention Based Classifier

Bowen Wang, Liangzhi Li, Jiahao Zhang, Yuta Nakashima, Hajime Nagahara

Experience

Education

Awards