Reexpress
官方为您的搜索、软件和数据科学工作流启用相似度-距离-幅度统计验证
你可以用 Reexpress MCP 做什么?
- 验证LLM响应 — 让您的助手对其自身答案运行
Reexpress,以获取针对OpenVerification1数据集的统计置信度估计。 - 更新验证模型 — 使用
ReexpressAddTrue或ReexpressAddFalse来纠正验证结果,并改进未来的概率估计。 - 控制文件访问 — 通过
ReexpressDirectorySet和ReexpressFileSet指定哪些本地文件被发送到LLM API。 - 生成HTML输出 — 请求生成验证结果的静态HTML报告,以便共享或审阅。
文档
Reexpress 模型上下文协议(MCP)服务器
适用于工具调用型大语言模型(例如 Claude Fable 5)以及运行在 macOS(Apple 芯片上的 Tahoe 26 或更高版本)或 Linux 上的 MCP 客户端
视频概览1:此处


Reexpress MCP 服务器是一款即插即用的解决方案,可为您的复杂大语言模型流水线以及日常在软件开发和数据科学场景中使用大语言模型进行搜索和问答时,添加最先进的统计验证能力。它是您 AI 工作流中首个可靠、统计稳健的 AI 第二意见。
只需安装 MCP 服务器,然后在聊天文本末尾添加 Reexpress 提示词。工具调用型大语言模型(例如 Anthropic 的大语言模型 Claude Fable 5)将使用提供的预训练 Reexpress 相似度-距离-幅度(SDM)估计器检查其响应,该估计器集成了 gpt-5.5-2026-04-23、gemini-3.1-pro-preview 和 gemini-embedding-2,并结合工具调用型大语言模型的输出,针对 OpenVerification1 数据集中的训练和校准示例数据库,计算预测不确定性的稳健估计。Reexpress 方法的独特之处在于,您可以轻松地将模型适配到您的任务:只需在验证完成后调用 ReexpressAddTrue 或 ReexpressAddFalse 工具,后续对 Reexpress 工具的调用将动态考虑您的更新来计算验证概率。我们还附带了模型的训练脚本,以便在需要更实质性的更改或希望使用替代底层大语言模型时,您可以运行完整的重新训练。
[!NOTE] 除了为您(用户)提供基于指令的输出置信度的原则性估计外,工具调用型大语言模型本身还可以利用验证输出逐步优化其答案,判断是否需要额外的外部资源或工具,或者是否已陷入僵局而需要向您寻求进一步澄清或信息。这就是我们所说的基于 SDM 验证的推理——这是 AI 工具集中一项全新的能力,我们认为它将为大语言模型和大语言模型智能体开辟更广泛的应用场景,无论是个人用户还是企业用户。
数据仅通过标准的 LLM API 调用发送至 Azure/OpenAI 和 Google,其中 gemini-3.1-pro-preview 调用通过 API 获得标准网络搜索访问权限;SDM 估计器的所有处理均在您本地计算机上完成。Reexpress MCP 采用简单、保守但有效的文件访问系统:您通过文件访问工具 ReexpressDirectorySet() 和 ReexpressFileSet() 明确指定要发送给 LLM API 的附加文件(如果有)。
版本 2.5.0 的新增内容
版本 2.5.0 实现了研究说明:嵌套相似度-距离-幅度估计器中描述的嵌套估计器。该方法简单地对类别条件和预测条件准确率运行递减概率阈值的校准算法。结果是,不再是单一区域,绝大多数校准点可以被分配到某个区域,该区域的类别条件和预测条件准确率估计至少大于 0.5。未分配到任何区域的剩余点可被视为有效的分布外点。最保守的区域保留其原有的解释和行为,我们发现嵌套区域为裁决剩余点的相对概率提供了有意义的排序。
此外,代码库已精简,移除了语言模型后训练代码。将单独发布一个仓库,用于微调网络的底层权重。
已发布模型在其他方面与版本 2.4.x 相同。它使用相同的数据进行校准,并以 gpt-5.5-2026-04-23 和 gemini-3.1-pro-preview 作为生成模型。详情请参阅版本 2.4.0 模型卡。
其他说明请参阅 changelog.md。
系统要求
MCP 服务器可在 Linux 和 macOS 上运行。主要要求是运行 MCP 服务器的机器需要能够在本地运行一个小的 300 万参数 PyTorch 模型,因此计算要求极低。(正如所写:只有 300 万百万参数;不是 30 亿十亿参数。该模型由 gemini-embedding-2 上的 SDM 激活以及两个 API 语言模型的分类输出组成。)
安装
请参阅 INSTALL.md。
[!TIP] 与其他 MCP 服务器相比,Reexpress MCP 服务器的设置相对简单,但我们假定您对 LLM、MCP 和命令行工具有所了解。我们的目标受众是开发人员和数据科学家。只添加来自您信任来源的其他 MCP 服务器,并注意其他 MCP 工具可能以意外方式改变我们 MCP 服务器的行为。
配置选项
请参阅 CONFIG.md。
使用方法
请参阅 documentation/HOW_TO_USE.md。
使用工具调用的输出生成静态 HTML
请参阅 documentation/OUTPUT_HTML.md。
指南
请参阅 documentation/GUIDELINES.md。
常见问题解答
请参阅 documentation/FAQ.md。
训练和校准数据
在 OpenVerification1 上的评估
系统演示论文
我们的系统演示论文《可内省、可更新且具有不确定性感知的语言模型指令遵循分类》的副本(特别关注 Reexpress MCP 服务器版本 2.1.0)包含在此处。用于复现分析的支持脚本包含在此处。
版本 2.4.0 的模型卡(重点介绍自系统演示论文以来的更改)可在此处获取。

引用
如果您觉得此软件有用,请考虑引用以下同行评审论文:
@inproceedings{Schmaltz-2026-SimilarityDistanceMagnitudeActivations,
title = "Similarity-Distance-Magnitude Activations",
author = "Schmaltz, Allen",
editor = "Liakata, Maria and
Moreira, Viviane P. and
Zhang, Jiajun and
Jurgens, David",
booktitle = "Findings of the {A}ssociation for {C}omputational {L}inguistics: {ACL} 2026",
month = jul,
year = "2026",
address = "San Diego, California, United States",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2026.findings-acl.1109/",
doi = "10.18653/v1/2026.findings-acl.1109",
pages = "22037--22057",
ISBN = "979-8-89176-395-1",
abstract = "We introduce the Similarity-Distance-Magnitude (SDM) activation function, a more robust and interpretable formulation of the standard softmax activation function, adding Similarity (i.e., correctly predicted depth-matches into training) awareness and Distance-to-training-distribution awareness to the existing output Magnitude (i.e., decision-boundary) awareness, and enabling interpretability-by-exemplar via dense matching. We further introduce the SDM estimator, based on a data-driven partitioning of the class-wise empirical CDFs via the SDM activation, to control the class- and prediction-conditional accuracy among selective classifications. When used as the final-layer activation over pre-trained language models for selective classification, the SDM estimator is more robust to covariate shifts and out-of-distribution inputs than existing calibration methods using softmax activations, while remaining informative over in-distribution data."
}
@inproceedings{Schmaltz-2026-ReexpressMCPServer,
author = {Schmaltz, Allen},
title = {Introspectable, Updatable, and Uncertainty-aware Classification of Language Model Instruction-following},
year = {2026},
isbn = {9798400724152},
publisher = {Association for Computing Machinery},
address = {New York, NY, USA},
url = {https://doi.org/10.1145/3786335.3813214},
doi = {10.1145/3786335.3813214},
abstract = {In this system demonstration paper, we introduce an open-source implementation for training and testing Similarity-Distance-Magnitude (SDM) estimators for the task of binary classification of instruction-following of closed-weight language models (LMs). This SDM estimator provides an approximately conditional estimate of the predictive uncertainty over instruction-following, conditional on multiple closed-weight LMs and the representation space of an open-weight model. While it would be more robust to use as input to the SDM estimator the hidden-states of the underlying models, this indirect, compositional proxy is more reliable than verbalized uncertainty and adds a means of auditing the predictions against data with known labels. We release the code as an MCP Server to simplify adding interpretability-by-exemplar and locally updatable, uncertainty-aware instruction-following to agent-based pipelines. We further release OpenVerification1, a balanced set of over two million examples of instruction-following and associated rationales from recent closed-weight LMs, for bootstrapping domain-specific estimators. Finally, we discuss limitations of estimating the predictive uncertainty without access to the hidden-states of the tool-calling LM and provide practical guidance for applications.},
booktitle = {Proceedings of the ACM Conference on AI and Agentic Systems},
pages = {1259–1269},
numpages = {11},
keywords = {Approximately conditional calibration, Interpretability-by-exemplar, Classification of instruction-following, Model ensembles},
location = {
},
series = {CAIS '26}
}
Footnotes
-
The 输出格式自视频中使用的 v1.0.0 以来已更改。请参阅 changelog.md。 ↩
