Back to Main Conference 2026
LREC 2026main

Building Effective Japanese Medical LLMs with an Open Recipe for Domain Adaptation through Continued Pre-training

Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026)

DOI:10.63317/47uvbxqph5ph

Abstract

In high-stakes domains such as medicine, ensuring transparency of the training corpus is essential, with careful consideration of local healthcare landscapes; however, the majority of existing medical large language models (LLMs) have not disclosed the details of their training corpora. Here, we introduce an open recipe for domain adaptation of LLMs to the Japanese medical domain. We employed fully open-source Japanese general-domain LLMs as base models, whose pre-training datasets are also disclosed. To establish effective corpora for domain adaptation through continued pre-training, we started with small-scale medical datasets and ultimately constructed a medical corpus consisting of 79.6B tokens, incorporating local clinical guidelines, medical textbooks, and other domain-specific resources. The resulting LLM from continued pre-training, namely SIP-med-llm-8x13B, with an active parameter count of 22B, demonstrated favorable accuracy on benchmarks including the Japanese National Medical Examination. This performance was comparable to that of 70B-parameter open-weight models whose construction details remain non-transparent. This represents the first case in the Japanese medical field where complete corpus details have been disclosed for fully from-scratch development, providing important insights for future efforts to construct medical LLMs tailored to the specific characteristics of local contexts. The model is available publicly at this Hugging Face repository: https://huggingface.co/SIP-med-LLM/SIP-jmed-llm-2-8x13b-OP-instruct.

Details

Paper ID
lrec2026-main-817
Pages
pp. 10405-10423
BibKey
aizawa-etal-2026-building
Editor
N/A
Publisher
European Language Resources Association (ELRA)
ISSN
2522-2686
ISBN
978-2-493814-49-4
Conference
The Fifteenth Language Resources and Evaluation Conference (LREC 2026)
Location
Palma, Mallorca, Spain
Date
11 May 2026 16 May 2026

Authors

  • AA

    Akiko Aizawa

  • YA

    Yuki Arase

  • FC

    Fei Cheng

  • JH

    Jiahao Huang

  • ZH

    Zhiyi Huang

  • JJ

    Junfeng Jiang

  • TK

    Teruhito Kanazawa

  • DK

    Daisuke Kawahara

  • KK

    Kazuma Kobayashi

  • TK

    Takashi Kodama

  • SK

    Sadao Kurohashi

  • YO

    Yusuke Oda

  • YT

    Yuma Tsuta

  • ZW

    Zhen Wan

  • ZY

    Zhishen Yang

  • RY

    Rio Yokota

Links