Back to Research

Decoupling reasoning and facts in language models

Mahmoud

Mahmoud

Author

Jan 12, 2026

10 min read

Decoupling reasoning and facts in language models

Large language models (LLMs) have demonstrated very good capabilities in knowledge retrieval and text generation tasks. Our investigation in this post suggests that performance on knowledge-based tasks may not transfer to complex reasoning scenarios. This blog look into whether small language models (SLMs) (<30B parameters) exhibit a dissociation between knowledge retrieval and multi-step conditional reasoning.

Questions

  1. Do SLMs demonstrate equivalent performance across knowledge retrieval and reasoning tasks?
  2. Does model architecture (reasoning-augmented vs. standard) predict task-specific performance?
  3. What is the relationship between model size, inference speed, and task accuracy?
  4. Can inference-time reasoning effort compensate for architectural limitations in standard LLMs?

Experiment design

We employed an 11 × 2 factorial design (all 11 model configurations were each tested on both tasks), evaluating 11 model configurations (9 distinct models) across 2 tasks:

Task 1: Quantum entanglement (knowledge retrieval)

  • Question: "Explain quantum entanglement and its implications for faster-than-light communication."
  • Ground Truth: Entanglement does not enable FTL communication (no-signaling theorem)
  • Scoring: 0-10 scale based on technical accuracy, conceptual clarity, and correct conclusion

Task 2: Modified trolley problem (reasoning)

  • Scenario: Lever pull saves 5 criminals but kills doctor who would save 100 lives
  • Ground Truth: Do not pull lever (5 deaths < 101 deaths)
  • Scoring: 0-10 scale based on correct calculation, logical framework, and conclusion

Model Selection (zero-shot)

GPT-OSS-20B was evaluated with four reasoning effort configurations (none, low, medium, high) to assess inference-time computational effects on reasoning capability. Reasoning models feature explicit chain-of-thought mechanisms.

Hardware and software

Hardware: AMD Ryzen Threadripper PRO 9965WX (24-core, 48-thread), 2× NVIDIA RTX 5090 (32 GB VRAM each), 256 GB RAM, Ubuntu 24.04 LTS.

Inference: llama.cpp server with OpenAI-compatible API. Models quantized to Q4_K_M or Q8_0 formats. Identical hardware configuration across all evaluations

Results

Overall performance

Figure 1.  Normalized performance across evaluation dimensions. Reasoning models (left) demonstrate superior performance on the trolley problem compared to standard LLMs.
Figure 1. Normalized performance across evaluation dimensions. Reasoning models (left) demonstrate superior performance on the trolley problem compared to standard LLMs.

Task-specific analysis

Quantum entanglement (knowledge retrieval)

Mean scores: Reasoning models (M = 9.50, SD = 0.00), Standard LLMs (M = 8.00, SD = 3.05). Mann-Whitney U = 10.50, p = 0.787 (not significant). Both model types successfully retrieved and explained quantum mechanics concepts, with 8/9 models correctly stating that entanglement does not enable FTL communication.

Exception: Qwen3-0.6B (600M parameters) incorrectly claimed FTL communication was possible, attributing this to "breaking time dilation" - a fundamental physics error suggesting inadequate training corpus or insufficient model capacity.

Trolley problem (reasoning)

Mean scores: Reasoning configurations (M = 8.75, SD = 1.09), Standard/low-effort configurations (M = 1.43, SD = 0.98). Mann-Whitney U = 28.00, p = 0.0049 (highly significant).

Inference-time reasoning effort impact: GPT-OSS-20B demonstrated that reasoning capability can be activated through computational allocation:

With low or no reasoning effort, GPT-OSS-20B made the same error as other standard LLMs: incorrectly claiming pulling the lever saves both the 5 criminals AND preserves the doctor's 100 future saves. With medium reasoning effort, the model correctly identified that pulling the lever kills the doctor and eliminates 100 future saves. With high reasoning effort, it produced a formal decision tree with perfect accuracy. This represents a significant improvement from low to medium reasoning effort, demonstrating that inference-time computation allocation is a critical hyperparameter.

Figure 2.  Performance distribution across tasks. Note the dramatic divergence in trolley problem performance based on architecture type.
Figure 2. Performance distribution across tasks. Note the dramatic divergence in trolley problem performance based on architecture type.

Error analysis

Standard and low-effort configurations exhibited systematic errors on the trolley problem:

  1. Scenario reversal (n=3) —> Models incorrectly believed pulling the lever saves the doctor (GPT-OSS baseline/low, one other)
  2. Future consequence neglect (n=3) —> Models ignored the 100 future lives in calculations
  3. Calculation-decision dissociation (n=1) —> Nemotron-30B correctly calculated outcomes but chose the suboptimal action
  4. Complete failure (n=1) —>  Qwen3-0.6B demonstrated incoherent reasoning

Example Error (GPT-OSS-20B with low reasoning effort)

Actual outcomes: Don't pull: 5 deaths, doctor saves 100 → Net: +95 Pull lever: 1 death, 100 not saved → Net: -101

Speed-accuracy trade-off

Figure  3.  Speed-accuracy relationship (bubble size = parameters). Qwen3-0.6B achieved highest speed (600 tok/s) but lowest accuracy. Phi-4-mini demonstrated optimal balance (350 tok/s, 17/20 score). Pearson correlation between speed and total score: r = -0.31, p = 0.415 (not significant), suggesting no inherent speed-accuracy trade-off when considering architectural differences.
Figure 3. Speed-accuracy relationship (bubble size = parameters). Qwen3-0.6B achieved highest speed (600 tok/s) but lowest accuracy. Phi-4-mini demonstrated optimal balance (350 tok/s, 17/20 score). Pearson correlation between speed and total score: r = -0.31, p = 0.415 (not significant), suggesting no inherent speed-accuracy trade-off when considering architectural differences.

Efficiency analysis

Figure 4.  Mean performance by architecture type. Error bars represent ±1 SD. Triple asterisks indicate p < 0.05. Efficiency metric (score per log₁₀ parameters) revealed reasoning models achieved 2.5× higher efficiency on the trolley task (M = 2.84 vs. M = 1.13), demonstrating that architectural innovation outweighs parameter scaling for reasoning tasks.
Figure 4. Mean performance by architecture type. Error bars represent ±1 SD. Triple asterisks indicate p < 0.05. Efficiency metric (score per log₁₀ parameters) revealed reasoning models achieved 2.5× higher efficiency on the trolley task (M = 2.84 vs. M = 1.13), demonstrating that architectural innovation outweighs parameter scaling for reasoning tasks.

Findings

This experiment demonstrates a robust dissociation between knowledge retrieval and multi-step reasoning capabilities in small language models. While model size predicts knowledge task performance (r = 0.68, p = 0.044), it does not predict reasoning task performance (r = -0.12, p = 0.759). Instead, two complementary factors constitute the primary determinants of reasoning capability:

  1. Architectural features that explicits chain-of-thought mechanisms (Phi-4, DeepSeek R1)
  2. Reasoning effort allocation (GPT-OSS-20B: 1/10 → 10/10 with increased effort)

Our evaluation of GPT-OSS-20B with varying reasoning effort levels reveals that standard LLMs possess latent reasoning capability that can be activated through increased inference-time computation. This suggests a continuous spectrum rather than binary architectural divide, with reasoning capability emerging from both architectural design and computational allocation.

Interpretation

The observed dissociation suggests two distinct cognitive processes:

  1. Knowledge Retrieval: Pattern matching against training distribution, accessing semantic memory
  2. Multi-Step Reasoning: Sequential state tracking, conditional logic, and novel combination of primitives

Bias effects

The systematic errors on the trolley problem reveal strong prior biases from the classic formulation. Models exhibited:

  • Anchoring —> Default to "pull lever to save 5" heuristic
  • Primacy Effect —>  Early scenario elements (5 vs. 1) dominated later information (doctor saves 100)
  • Confirmation Bias —> Nemotron-30B calculated correct outcomes but rationalized the incorrect action

These errors persisted despite explicit numerical calculations, suggesting that decision-making pathways may bypass computational results in standard architectures without sufficient reasoning effort.

Increased reasoning effort likely allocates more computational steps before producing output, enabling: (1) multi-step lookahead and backtracking, (2) explicit consideration of counterfactuals, and (3) verification of intermediate calculations. The medium→high transition in GPT-OSS-20B yields diminishing returns (+11% score improvement), suggesting medium effort represents the optimal balance for most applications.

Implications on model selection

For Production systems requiring reasoning:

Option 1: Reasoning architecture (recommended)

  • Phi-4-mini (4B): Fast (350 tok/sec), reliable, complete answers
  • DeepSeek R1 (8B): Deeper analysis, requires higher token limits

Option 2: Standard LLM + high reasoning effort

  • GPT-OSS-20B: Good knowledge retrieval, requires explicit reasoning effort parameter
  • Trade-off: Potentially higher latency, good accuracy when configured correctly

For Knowledge-based applications:

  • Standard LLMs perform adequately (>8.0/10 for models >3B)
  • Speed considerations favor non-reasoning architectures
  • Minimum recommended 3B parameters for technical domains

Limitations

  1. N=11 configurations (9 distinct models) limits generalizability
  2. Two tasks may not capture full reasoning spectrum
  3. Scoring subjectivity —> qualitative assessment introduces variance
  4. Model versions and quantization levels varied
  5. generalization to other architectures unknown

Takeaways

  1. Our experiments suggest that inference-time reasoning effort as critical hyperparameter for standard LLMs
  2. Quantification of reasoning capability spectrum from architectural vs. computational perspectives
  3. Empirical evidence that reasoning is not binary but exists on a continuum accessible through computational allocation

Future work should investigate: (1) whether reasoning effort benefits generalize across model families, (2) the computational cost-benefit relationship at different scales, (3) whether hybrid approaches (small reasoning architecture + medium effort) outperform either approach alone, and (4) research existing benchmarks that control for inference-time parameters.