AI駆動科学では再現性が鍵であることを実証(Study finds reproducibility is key to AI-driven science)

2026-08-03 スタンフォード大学

スタンフォード大学は、AIの急速な普及が科学研究の再現性(reproducibility)に与える影響について開催したCORES(Center for Open and REproducible Science)シンポジウムの議論を紹介した。研究者らは、AIは文献調査、データ解析、コード作成、統計解析などを効率化し、研究の透明性や再現性を高める可能性がある一方、AIが生成した結果を無批判に利用すると誤りやバイアスが拡散し、再現性を損なう危険性もあると指摘した。そのため、AIを研究の代替ではなく支援ツールとして位置付け、使用したモデルやプロンプト、データ、解析手順を詳細に記録・公開するとともに、人間による検証とオープンサイエンスの実践が不可欠であると強調した。AI時代には、研究成果だけでなく研究プロセス全体の透明性を確保することが、信頼できる科学の基盤になると結論付けている。

AI駆動科学では再現性が鍵であることを実証(Study finds reproducibility is key to AI-driven science)
A study conducted at four laboratories across the country demonstrated how experimental protocol and equipment standardization govern result variability. Recognizing this variability is essential when incorporating real-world data into AI/machine learning models. | Adam Hoffman and Greg Stewart / SLAC National Accelerator Laboratory

<関連情報>

データ駆動型モデリングのためのラウンドロビン試験によるCO2水素化反応中の触媒活性および失活における不確実性の定量化 Quantifying uncertainty in catalyst activity and deactivation during CO2 hydrogenation via round-robin testing for data-driven modelling

Selin Bac,Dongjae Shin,Seunghwa Hong,Jake Heinlein,Anastassiya Khan,Greg Barber,Zhihengyu Chen,Michael M. Albrechtsen,Christopher Tassone,Robert M. Rioux,Matteo Cargnello,Simon R. Bare,Kirsten Winther,Phillip Christopher & Adam S. Hoffman
Nature Catalysis  Published:31 July 2026
DOI:https://doi.org/10.1038/s41929-026-01559-y

Abstract

Machine learning (ML) is rapidly emerging as a catalyst discovery method, whose success depends on curated experimental datasets with defined uncertainties. Despite this need, uncertainty in catalyst performance is rarely quantified across datasets generated using multiple reactors. Here we present a four-laboratory round-robin study of Rh/TiO2 catalysts for CO2 hydrogenation that demonstrates that accounting for both intra- and interlaboratory variability is essential for identifying features for experimentally derived ML models. Even with identical catalyst batches and testing protocols, relationships between inputs (reaction temperature, Rh loading and synthesis method) and outputs (conversion, selectivity and CO and CH4 production rates) that were clear in intralaboratory studies became statistically insignificant when interlaboratory variability was included. Heat management emerged as a key contributor to this variability. Our work demonstrates that uncertainty analysis must be included in the selection of performance metrics and input features for ML models while revealing sources of variability that limit rigour and reproducibility.

1603情報システム・データ工学
ad
ad
Follow
ad
タイトルとURLをコピーしました