Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Forethought: Verifiable Reasoning from Neurosymbolic Primitive Programming

About

Current agentic workflows usually involve decomposing user requests into sequences of tool calls with correctly resolved parameters, the results of which are processed through reasoning traces in the language model's context window. The prevailing route to improve such reasoning is test-time scaling, which trains models to search over long chains of thought; but the resulting capability is entangled in model weights, is not verifiable step-by-step, and is costly at inference. We present Forethought, a neurosymbolic reasoning system that instead treats reasoning as an explicit, verifiable program, that builds from a library of symbolic and neural primitives which are composed through a domain-specific language. The result are reasoning programs, which are concrete representations of the model's work, and as such can be inspected and modified before deployment. Instantiated as a tool-calling execution kernel and evaluated across five benchmarks, Forethought improves base-model accuracy by about 30% relative and outperforms vanilla prompting, reinforcement learning scaffolds, and prompt-evolution methods, enabling small models to match or exceed frontier models capabilities. In a direct comparison, a non-reasoning model augmented with Forethought competes with a dedicated reasoning model while requiring roughly three orders of magnitude less post-training investment, and remains model-agnostic and auditable.

Vishvesh Bhat, Jay Vaghasiya, Emmanuel Anaya Gonzalez• 2026

Related benchmarks

TaskDatasetResultRank
Tool UseAggregate Performance
Avg1 Score76.8
20
Tool UseBFCL
BFCL v4 Score73
20
Tool UseLiveMCPBench
LiveMCP Score76
20
Showing 3 of 3 rows

Other info

Follow for update