Reexpress

官方

為您的搜尋、軟體及資料科學工作流程啟用相似度-距離-幅度統計驗證

你可以用 Reexpress MCP 做什麼?

  • 驗證 LLM 回應 — 要求您的助理對其自身答案執行 Reexpress,以取得針對 OpenVerification1 資料集的統計信賴度估計。
  • 更新驗證模型 — 使用 ReexpressAddTrue 或 ReexpressAddFalse 來修正驗證結果,並改善未來的機率估計。
  • 控制檔案存取 — 透過 ReexpressDirectorySet 和 ReexpressFileSet 指定哪些本機檔案會傳送至 LLM API。
  • 產生 HTML 輸出 — 要求產生驗證結果的靜態 HTML 報告,以供分享或檢閱。

文件

Reexpress Model-Context-Protocol (MCP) 伺服器

適用於工具呼叫型 LLM(例如 Claude Fable 5)以及執行於 macOS(Apple 晶片上的 Tahoe 26 或更新版本)或 Linux 的 MCP 用戶端

影片總覽1:點此

Watch the YouTube video

Screenshot image of the rendered HTML output from the Reexpress tool.

Re

Reexpress MCP Server 是一個可直接整合的解決方案,可為您的複雜 LLM 管線,以及您在軟體開發與資料科學情境中日常使用 LLM 進行搜尋和問答時,加入最先進的統計驗證能力。這是第一個可靠、統計上穩健的 AI 第二意見系統,適用於您的 AI 工作流程。

只需安裝 MCP 伺服器,然後在聊天文字末尾加入 Reexpress 提示詞。工具呼叫型 LLM(例如 Anthropic 的 LLM 模型 Claude Fable 5)便會使用所提供的預訓練 Reexpress 相似度-距離-幅度(SDM)估計器來檢查其回應,該估計器整合了 gpt-5.5-2026-04-23、gemini-3.1-pro-preview 和 gemini-embedding-2,以及工具呼叫型 LLM 的輸出,並針對 OpenVerification1 資料集中的訓練與校正範例資料庫,計算預測不確定性的穩健估計值。Reexpress 方法的獨特之處在於,您可以輕鬆地將模型調整至您的任務:只需在驗證完成後呼叫 ReexpressAddTrue 或 ReexpressAddFalse 工具,未來對 Reexpress 工具的呼叫便會在計算驗證機率時動態納入您的更新。我們也附上了模型的訓練腳本,因此當需要更實質性的變更,或您想使用其他底層 LLM 時,您可以執行完整的重新訓練。

[!NOTE] 除了為您(使用者)提供根據指示對輸出信心的原則性估計之外,工具呼叫型 LLM 本身也可以使用驗證輸出逐步完善其答案、判斷是否需要額外的外部資源或工具,或判斷是否已陷入僵局而需要向您請求進一步的釐清或資訊。這就是我們所稱的以 SDM 驗證進行推理——這是 AI 工具組中一項全新的能力,我們相信它將為個人和企業開啟更廣泛的 LLM 與 LLM 代理使用案例。

資料僅透過標準的 LLM API 呼叫傳送至 Azure/OpenAI 和 Google,其中 gemini-3.1-pro-preview 的呼叫會透過 API 獲得標準的網路搜尋存取權限;SDM 估計器的所有處理都在您的電腦上本地執行。Reexpress MCP 具有簡單且保守但有效的檔案存取系統:您可以透過檔案存取工具 ReexpressDirectorySet() 和 ReexpressFileSet() 明確指定要傳送至 LLM API 的其他檔案(如果有的話)。

2.5.0 版的新功能

2.5.0 版實作了研究備註:巢狀相似度-距離-幅度估計器中所述的巢狀估計器。此方法簡單地針對類別條件與預測條件準確度,以降序的機率閾值執行校正演算法。其結果是,與其只有單一區域,絕大多數的校正點都可以被分配到一個區域,該區域的類別條件與預測條件準確度估計至少大於 0.5。其餘未分配到任何區域的點可視為有效的分布外(out-of-distribution)資料。最保守的區域保留其原有的解釋和行為,我們發現巢狀區域提供了有意義的排序,可用於裁決其餘點的相對機率。

此外,程式碼庫已精簡,移除了語言模型後訓練程式碼。一個獨立的儲存庫將發布,用於微調網路的底層權重。

發布的模型在其他方面與 2.4.x 版相同。它使用相同的資料,並以 gpt-5.5-2026-04-23 和 gemini-3.1-pro-preview 作為生成模型進行校正。詳細資訊請參閱 2.4.0 版模型卡。

其他說明請參閱 changelog.md。

系統需求

MCP 伺服器可在 Linux 和 macOS 上執行。主要需求是執行 MCP 伺服器的機器需要能夠在本地執行一個小型 300 萬參數的 PyTorch 模型,因此運算需求極低。(是的,您沒看錯:只有 300 萬個參數,不是 30 億個參數。該模型由 gemini-embedding-2 上的 SDM 激活以及兩個 API 語言模型的分類輸出組成。)

安裝

請參閱 INSTALL.md。

[!TIP] 相較於其他 MCP 伺服器,Reexpress MCP 伺服器的設定相對簡單,但我們假設您對 LLM、MCP 和命令列工具有一定程度的熟悉。我們的目標受眾是開發人員和資料科學家。請僅從您信任的來源新增其他 MCP 伺服器,並請記住,其他 MCP 工具可能會以非預期的方式改變我們 MCP 伺服器的行為。

設定選項

請參閱 CONFIG.md。

使用方式

請參閱 documentation/HOW_TO_USE.md。

使用工具呼叫的輸出產生靜態 HTML

請參閱 documentation/OUTPUT_HTML.md。

指南

請參閱 documentation/GUIDELINES.md。

常見問題(FAQ)

請參閱 documentation/FAQ.md。

訓練與校正資料

請參閱 documentation/DATA.md。

在 OpenVerification1 上的評估

請參閱 documentation/EVAL.md。

系統示範論文

我們的系統示範論文「Introspectable, Updatable, and Uncertainty-aware Classification of Language Model Instruction-following」的副本已包含在此處,該論文特別聚焦於 Reexpress MCP Server 的 2.1.0 版。用於重現分析的支援腳本已包含在此處。

2.4.0 版的模型卡(重點說明自系統示範論文以來的變更)可在此處取得。

CAIS 2026 system demonstration poster.

引用

如果您覺得此軟體有用,請考慮引用以下經同儕審查的論文:

@inproceedings{Schmaltz-2026-SimilarityDistanceMagnitudeActivations,
    title = "Similarity-Distance-Magnitude Activations",
    author = "Schmaltz, Allen",
    editor = "Liakata, Maria  and
      Moreira, Viviane P.  and
      Zhang, Jiajun  and
      Jurgens, David",
    booktitle = "Findings of the {A}ssociation for {C}omputational {L}inguistics: {ACL} 2026",
    month = jul,
    year = "2026",
    address = "San Diego, California, United States",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2026.findings-acl.1109/",
    doi = "10.18653/v1/2026.findings-acl.1109",
    pages = "22037--22057",
    ISBN = "979-8-89176-395-1",
    abstract = "We introduce the Similarity-Distance-Magnitude (SDM) activation function, a more robust and interpretable formulation of the standard softmax activation function, adding Similarity (i.e., correctly predicted depth-matches into training) awareness and Distance-to-training-distribution awareness to the existing output Magnitude (i.e., decision-boundary) awareness, and enabling interpretability-by-exemplar via dense matching. We further introduce the SDM estimator, based on a data-driven partitioning of the class-wise empirical CDFs via the SDM activation, to control the class- and prediction-conditional accuracy among selective classifications. When used as the final-layer activation over pre-trained language models for selective classification, the SDM estimator is more robust to covariate shifts and out-of-distribution inputs than existing calibration methods using softmax activations, while remaining informative over in-distribution data."
}
@inproceedings{Schmaltz-2026-ReexpressMCPServer,
    author = {Schmaltz, Allen},
    title = {Introspectable, Updatable, and Uncertainty-aware Classification of Language Model Instruction-following},
    year = {2026},
    isbn = {9798400724152},
    publisher = {Association for Computing Machinery},
    address = {New York, NY, USA},
    url = {https://doi.org/10.1145/3786335.3813214},
    doi = {10.1145/3786335.3813214},
    abstract = {In this system demonstration paper, we introduce an open-source implementation for training and testing Similarity-Distance-Magnitude (SDM) estimators for the task of binary classification of instruction-following of closed-weight language models (LMs). This SDM estimator provides an approximately conditional estimate of the predictive uncertainty over instruction-following, conditional on multiple closed-weight LMs and the representation space of an open-weight model. While it would be more robust to use as input to the SDM estimator the hidden-states of the underlying models, this indirect, compositional proxy is more reliable than verbalized uncertainty and adds a means of auditing the predictions against data with known labels. We release the code as an MCP Server to simplify adding interpretability-by-exemplar and locally updatable, uncertainty-aware instruction-following to agent-based pipelines. We further release OpenVerification1, a balanced set of over two million examples of instruction-following and associated rationales from recent closed-weight LMs, for bootstrapping domain-specific estimators. Finally, we discuss limitations of estimating the predictive uncertainty without access to the hidden-states of the tool-calling LM and provide practical guidance for applications.},
    booktitle = {Proceedings of the ACM Conference on AI and Agentic Systems},
    pages = {1259–1269},
    numpages = {11},
    keywords = {Approximately conditional calibration, Interpretability-by-exemplar, Classification of instruction-following, Model ensembles},
    location = {
    },
    series = {CAIS '26}
}

Footnotes

  1. The 輸出格式自影片中使用的 v1.0.0 以來已有所變更。請參閱 changelog.md。 ↩