データに潜む「複数の法則」を自動で発見する新手法を開発ー1つの答えに絞らない新しいシンボリック回帰手法「DRSR」ー

2026-08-26 東京科学大学

東京科学大学、サイバーエージェントの研究者らは、データに含まれる 複数の法則を異なる数式として同時に発見するシンボリック回帰手法「DRSR(Diversified Residual Symbolic Regression)」 を開発した。従来の手法はデータ全体の予測誤差を最小化し、最も適した1つの数式を選ぶことが多かったが、外れ値や異なる現象による複数の傾向が混在する場合、別の有力な関係を見落とす可能性があった。DRSRは予測精度だけでなく、各数式が「どのデータを説明し、どこに誤差を残すか」という残差パターンの違いを評価し、特徴の異なる数式候補を探索・保持する。人工データでは外れ値に強く、複数の法則が混在するデータからそれぞれの関係を抽出できたほか、天文学データでも恒星の質量―光度関係と整合する複数候補を発見した。AIが単一の「正解」を提示するのではなく、研究者が複数の仮説を比較・検証できるため、物理学、材料科学、生命科学などの科学的発見を支援する技術として期待される。

データに潜む「複数の法則」を自動で発見する新手法を開発ー1つの答えに絞らない新しいシンボリック回帰手法「DRSR」ー
図1. DRSRにより観測データから複数の数式候補を提示する概念図。数式およびデータは説明のための模式例。

<関連情報>

多様化残差シンボリック回帰 Diversified Residual Symbolic Regression

Koki Ikeda, Masahiro Nomura, Ryoki Hamano
GECCO ’26: Proceedings of the Genetic and Evolutionary Computation Conference  Published: 10 July 2026
DOI:https://doi.org/10.1145/3795095.3805087

Abstract

Symbolic regression (SR) aims to discover explicit mathematical expressions that explain observed data and is widely used in domains where interpretability is essential. Because interpretability requires expressions to reflect meaningful regularities, SR is sensitive to observations that deviate from the dominant relationship. Such irregular observations, or outliers, are common in real-world data and can hinder SR from identifying underlying regularities. Robust regression mitigates this by downweighting observations with large residuals. However, deciding which observations should be treated as outliers is often ambiguous and depends on user interpretation and domain knowledge, a perspective largely overlooked in existing SR studies. This motivates approaches that present multiple candidate expressions, allowing users to examine different residual patterns and choose expressions consistent with their expertise. We propose diversified residual symbolic regression (DRSR), which achieves high predictive accuracy while promoting diversity with respect to residual patterns based on the Quality-Diversity paradigm. DRSR collects multiple expressions that fit the data well but differ in how residuals are distributed, enabling post-search selection aligned with domain knowledge. On a synthetic mixture dataset, DRSR produces more diverse expressions than conventional SR while capturing multiple underlying relationships. On a real-world astronomical dataset, DRSR discovers multiple expressions consistent with known physical relationships.

1504数理・情報
ad
ad
Follow
ad
タイトルとURLをコピーしました