Reexpress

公式

検索、ソフトウェア、データサイエンスのワークフローに類似性・距離・大きさの統計的検証を導入します。

Reexpress MCPで何ができますか?

  • LLM応答の検証 — アシスタントに自身の回答に対してReexpressを実行させ、OpenVerification1データセットに対する統計的信頼度推定を取得します。
  • 検証モデルの更新 — ReexpressAddTrueまたはReexpressAddFalseを使用して検証結果を修正し、将来の確率推定を改善します。
  • ファイルアクセスの制御 — ReexpressDirectorySetおよびReexpressFileSetを介して、LLM APIに送信するローカルファイルを指定します。
  • HTML出力の生成 — 検証結果の静的HTMLレポートをリクエストして、共有やレビューに使用します。

ドキュメント

Reexpress Model-Context-Protocol (MCP) サーバー

ツール呼び出し対応LLM(例:Claude Fable 5)およびmacOS(Appleシリコン上のTahoe 26以降)またはLinuxで動作するMCPクライアント向け

ビデオ概要1: こちら

Watch the YouTube video

Screenshot image of the rendered HTML output from the Reexpress tool.

Re

Reexpress MCP Serverは、複雑なLLMパイプラインや、ソフトウェア開発およびデータサイエンス環境における検索・QAのためのLLMの日常的な利用に、最先端の統計的検証を追加するためのドロップインソリューションです。これは、AIワークフローに対する初の信頼性が高く、統計的に堅牢なAIセカンドオピニオンです。

MCPサーバーをインストールし、チャットテキストの最後にReexpressプロンプトを追加するだけです。ツール呼び出し対応LLM(例:AnthropicのLLMモデルClaude Fable 5)は、提供された事前学習済みReexpress Similarity-Distance-Magnitude (SDM) estimatorを使用して応答をチェックします。この推定器は、gpt-5.5-2026-04-23、gemini-3.1-pro-preview、gemini-embedding-2をアンサンブルし、ツール呼び出し対応LLMの出力とともに、OpenVerification1データセットからのトレーニングおよびキャリブレーション例のデータベースに対する予測不確実性の堅牢な推定値を計算します。Reexpressメソッドに固有の点として、タスクに合わせてモデルを簡単に適応させることができます。検証が完了した後にReexpressAddTrueまたはReexpressAddFalseツールを呼び出すだけで、その後のReexpressツールへの呼び出しは、検証確率の計算時に更新内容を動的に考慮します。また、モデルのトレーニングスクリプトも含まれているため、より実質的な変更が必要な場合や、代替の基盤LLMを使用したい場合に、完全な再トレーニングを実行できます。

[!NOTE] 指示に基づく出力に対する原則に基づいた信頼度の推定値をユーザーに提供することに加えて、ツール呼び出し対応LLM自体が検証出力を使用して、回答を段階的に洗練したり、追加の外部リソースやツールが必要かどうかを判断したり、行き詰まった場合にさらなる明確化や情報を求めることができます。これが私たちがSDM検証による推論と呼ぶものです。これはAIツールキットにおけるまったく新しい機能であり、個人と企業の両方にとって、LLMおよびLLMエージェントのより幅広いユースケースを切り開くものと考えています。

データは、標準のLLM API呼び出しを介してAzure/OpenAIおよびGoogleにのみ送信され、gemini-3.1-pro-preview呼び出しにはAPIを介した標準のWeb検索アクセスが与えられます。SDM推定器の処理はすべてローカルコンピューター上で実行されます。Reexpress MCPには、シンプルで保守的ですが効果的なファイルアクセスシステムがあります。ファイルアクセスツールReexpressDirectorySet()およびReexpressFileSet()を使用して、LLM APIに送信する追加ファイル(ある場合)を明示的に指定して制御します。

バージョン2.5.0の新機能

バージョン2.5.0リリースでは、Research Note: Nested Similarity-Distance-Magnitude Estimatorsで説明されているネストされた推定器を実装しています。このアプローチでは、クラス条件および予測条件の精度に対して、降順の確率しきい値でキャリブレーションアルゴリズムを実行するだけです。その結果、単一の領域ではなく、キャリブレーションポイントの大部分を、クラス条件および予測条件の精度が0.5より大きいと推定される領域に割り当てることができます。どの領域にも割り当てられなかった残りのポイントは、事実上、分布外と見なすことができます。最も保守的な領域は、以前と同じ解釈と動作を維持し、ネストされた領域は、残りのポイントの相対的な確率を判定するための有意義なランキングを提供することがわかりました。

さらに、コードベースは合理化され、言語モデルのポストトレーニングコードは削除されました。ネットワークの基盤となる重みのファインチューニング用に、別のリポジトリがリリースされる予定です。

リリースされたモデルは、それ以外はバージョン2.4.xと同じです。同じデータを使用し、生成モデルとしてgpt-5.5-2026-04-23およびgemini-3.1-pro-previewを使用してキャリブレーションされています。詳細については、Version 2.4.0 model cardを参照してください。

追加の注意事項はchangelog.mdにあります。

システム要件

MCPサーバーはLinuxおよびmacOSで動作します。主な要件は、MCPサーバーを実行するマシンが、小さな300万パラメータのPyTorchモデルをローカルで実行できることです。そのため、計算要件は最小限です。(その通りです。30億パラメータではなく、300万パラメータのみです。モデルは、gemini-embedding-2上のSDMアクティベーションと、2つのAPI言語モデルの分類出力で構成されています。)

インストール

INSTALL.mdを参照してください。

[!TIP] Reexpress MCPサーバーは、他のMCPサーバーと比較してセットアップは簡単ですが、LLM、MCP、およびコマンドラインツールにある程度精通していることを前提としています。対象読者は開発者とデータサイエンティストです。信頼できるソースからの他のMCPサーバーのみを追加し、他のMCPツールが当社のMCPサーバーの動作を予期しない方法で変更する可能性があることに注意してください。

設定オプション

CONFIG.mdを参照してください。

使用方法

documentation/HOW_TO_USE.mdを参照してください。

ツール呼び出しの出力を使用した静的HTMLの生成

documentation/OUTPUT_HTML.mdを参照してください。

ガイドライン

documentation/GUIDELINES.mdを参照してください。

FAQ

documentation/FAQ.mdを参照してください。

トレーニングおよびキャリブレーションデータ

documentation/DATA.mdを参照してください。

OpenVerification1での評価

documentation/EVAL.mdを参照してください。

システムデモンストレーションペーパー

システムデモンストレーションペーパー「Introspectable, Updatable, and Uncertainty-aware Classification of Language Model Instruction-following」のコピーは、特にReexpress MCP Serverのバージョン2.1.0に焦点を当てたもので、こちらに含まれています。分析を再現するためのサポートスクリプトはこちらに含まれています。

システムデモンストレーションペーパー以降の変更点を強調したバージョン2.4.0のモデルカードは、こちらで入手できます。

CAIS 2026 system demonstration poster.

引用

このソフトウェアが役立つと思われた場合は、以下の査読付き論文を引用することを検討してください:

@inproceedings{Schmaltz-2026-SimilarityDistanceMagnitudeActivations,
    title = "Similarity-Distance-Magnitude Activations",
    author = "Schmaltz, Allen",
    editor = "Liakata, Maria  and
      Moreira, Viviane P.  and
      Zhang, Jiajun  and
      Jurgens, David",
    booktitle = "Findings of the {A}ssociation for {C}omputational {L}inguistics: {ACL} 2026",
    month = jul,
    year = "2026",
    address = "San Diego, California, United States",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2026.findings-acl.1109/",
    doi = "10.18653/v1/2026.findings-acl.1109",
    pages = "22037--22057",
    ISBN = "979-8-89176-395-1",
    abstract = "We introduce the Similarity-Distance-Magnitude (SDM) activation function, a more robust and interpretable formulation of the standard softmax activation function, adding Similarity (i.e., correctly predicted depth-matches into training) awareness and Distance-to-training-distribution awareness to the existing output Magnitude (i.e., decision-boundary) awareness, and enabling interpretability-by-exemplar via dense matching. We further introduce the SDM estimator, based on a data-driven partitioning of the class-wise empirical CDFs via the SDM activation, to control the class- and prediction-conditional accuracy among selective classifications. When used as the final-layer activation over pre-trained language models for selective classification, the SDM estimator is more robust to covariate shifts and out-of-distribution inputs than existing calibration methods using softmax activations, while remaining informative over in-distribution data."
}
@inproceedings{Schmaltz-2026-ReexpressMCPServer,
    author = {Schmaltz, Allen},
    title = {Introspectable, Updatable, and Uncertainty-aware Classification of Language Model Instruction-following},
    year = {2026},
    isbn = {9798400724152},
    publisher = {Association for Computing Machinery},
    address = {New York, NY, USA},
    url = {https://doi.org/10.1145/3786335.3813214},
    doi = {10.1145/3786335.3813214},
    abstract = {In this system demonstration paper, we introduce an open-source implementation for training and testing Similarity-Distance-Magnitude (SDM) estimators for the task of binary classification of instruction-following of closed-weight language models (LMs). This SDM estimator provides an approximately conditional estimate of the predictive uncertainty over instruction-following, conditional on multiple closed-weight LMs and the representation space of an open-weight model. While it would be more robust to use as input to the SDM estimator the hidden-states of the underlying models, this indirect, compositional proxy is more reliable than verbalized uncertainty and adds a means of auditing the predictions against data with known labels. We release the code as an MCP Server to simplify adding interpretability-by-exemplar and locally updatable, uncertainty-aware instruction-following to agent-based pipelines. We further release OpenVerification1, a balanced set of over two million examples of instruction-following and associated rationales from recent closed-weight LMs, for bootstrapping domain-specific estimators. Finally, we discuss limitations of estimating the predictive uncertainty without access to the hidden-states of the tool-calling LM and provide practical guidance for applications.},
    booktitle = {Proceedings of the ACM Conference on AI and Agentic Systems},
    pages = {1259–1269},
    numpages = {11},
    keywords = {Approximately conditional calibration, Interpretability-by-exemplar, Classification of instruction-following, Model ensembles},
    location = {
    },
    series = {CAIS '26}
}

Footnotes

  1. The 出力形式は、ビデオで使用されているv1.0.0以降に変更されています。changelog.mdを参照してください。 ↩