Back to Research

The Limits of prompt optimization in high-alignment model pairs

Eya Gammoudi

Eya Gammoudi

Author

Dec 25, 2025

5 min read

The Limits of prompt optimization in high-alignment model pairs

In a recent experiment, we migrated two agents “Find Locations” and “Assessment” from an AI SDK–based implementation to DSPy.

In this case study, two outcomes stood out:

  • A lower-capacity execution model (gemini-2.0-flash-lite) reproduced the structured-output behavior of a higher-capability reference model (gemini-2.5-flash) with near-identical accuracy on the evaluated tasks.
  • DSPy’s Refine algorithm, despite running multiple optimization rounds, introduced only minimal prompt changes and yielded no measurable improvement in output quality.

These observations raise broader technical questions about model capacity, implicit generalization, and the practical limits of automated prompt optimization in low-entropy, structured tasks:

  • Implicit generalization

    To what extent do modern instruction-tuned models internalize structural reasoning patterns during pre-training and alignment, reducing reliance on explicit prompt engineering?

  • Limits of prompt optimization

    When a model already generalizes correctly from minimal prompting, what headroom remains for refinement algorithms such as DSPy’s Refine?

    Is there an effective upper bound on prompt-level optimization when reference and execution model output distributions are already closely aligned?

These questions motivated a deeper examination of how DSPy Refine behaves when baseline model behavior is already stable and near-optimal, rather than error-prone.

The experiment

In our experiment, teacher–student is used purely in the experimental sense, common in prompt optimization and imitation learning:

  • Teacher: a stronger model used to generate reference outputs (gemini-2.5-flash)
  • Student: a cheaper/faster model optimized to match those outputs (gemini-2.0-flash-lite)
python
import dspy
from dspy.optimize import Refine

#  Teacher and Student Models
teacher = dspy.LM(
    model="gemini-2.5-flash", 
    temperature=0.0
)

student = dspy.LM(
    model="gemini-2.0-flash-lite",  
    temperature=0.0
)

dspy.configure(lm=student)

#  DSPy Signatures
class FindLocationsSig(dspy.Signature):
    query = dspy.InputField()
    result = dspy.OutputField(desc="Structured JSON location data")

class AssessmentSig(dspy.Signature):
    details = dspy.InputField()
    assessment = dspy.OutputField(desc="Structured JSON evaluation")

FindLocations = dspy.ChainOfThought(FindLocationsSig)
Assessment   = dspy.ChainOfThought(AssessmentSig)

#  Teacher Output Generation
def get_teacher_output(signature, **kwargs):
    """
    Generate high-quality teacher baselines.
    Parameters:
    signature : dspy.Signature or dspy.ChainOfThought
        The DSPy module representing the task to run (e.g:FindLocations or Assessment).
        Calling `signature(**kwargs)` executes the module with the provided inputs.
    """
    dspy.configure(lm=teacher)
    out = signature(**kwargs)
    dspy.configure(lm=student)
    return out

teacher_samples = [
    {
        "input": {"query": "Find one EV charging startion in Berlin"},
        "target": get_teacher_output(FindLocations, query="EV charging startion in Berlin")
    },
    {
        "input": {"details": "New EV charging station in Stuttgart"},
        "target": get_teacher_output(Assessment, details="New EV charging station in Stuttgart")
    }
]

# DSPy Refine
refiner = Refine(
    metric=lambda gold, pred: dspy.evaluate.eq(gold, pred),
    N=5,
    threshold=1.0  # Only accept perfect matches from the teacher
)

refined_FindLocations = refiner(FindLocations, teacher_samples)
refined_Assessment    = refiner(Assessment, teacher_samples)

Two production agents were fully migrated from the AI SDK to DSPy:

  • “Find Locations Agent” → Extracts structured location data
  • “Assessment Agent” → Produces structured assessments based on inputs

The pipeline followed a classic teacher–student learning paradigm:

  1. Teacher model outputs generated as reference (high-quality baselines)
  2. Student model wrapped in DSPy (LLM) as the execution engine
  3. DSPy’s Refine applied to optimize prompts against teacher behavior
  4. Evaluations compared:
    • Output vs Output (teacher vs student)
    • Prompt vs Prompt (original vs refined)
    Article image

The results

Across both agents, two major outcomes emerged:

1. The student model almost matched the teacher model.

The student model produced:

  • Nearly identical structured outputs
  • Minimal differences in phrasing or formatting
  • No degradation in correctness or task execution

This indicates that for these narrowly defined, structured tasks, model capacity was not the bottleneck. The student’s instruction-following capabilities were already strong enough to replicate the teacher’s behavior with very little drift.

2. DSPy Refine introduced extremely small changes to prompts.

Refine behaved conservatively:

  • It added tiny clarifications, often in the form of short examples
  • It did not restructure, rewrite, or meaningfully alter the base prompt
  • The refined prompt produced outputs effectively identical to those from the original prompt

Example:

Before refinement:

Below is an excerpt from the prompt used in our experiment:

plain text
DATA_COLLECTION_PROMPT = """
Role: You are a specialized Location Data Collection Agent focused exclusively on gathering
comprehensive location data for EV charging infrastructure deployment analysis. Your role is
data collection ONLY - you do not perform analysis or generate reports.
**CRITICAL: ALWAYS respond in the same language as the user's input/query.**

**PRIMARY OBJECTIVE**: Collect all necessary raw data about a location and store it in session
state for use by downstream analysis and reporting agents.

....

**OUTPUT**:

Your final response should be a raw json containing all the following tags:
- `location_coordinates`: {"lat": float, "lng": float, "formatted_address": string}
- `search_categories`: [list of extracted category strings]
- `nearby_places`: {category: {"count": int, "places": [...]}}
- `existing_charging_stations`: [list of charging station data]
- `traffic_patterns`: {traffic and route analysis data}

After refinement:

plain text
DATA_COLLECTION_PROMPT = """
Role: You are a specialized Location Data Collection Agent focused exclusively on gathering
comprehensive location data for EV charging infrastructure deployment analysis. Your role is
data collection ONLY - you do not perform analysis or generate reports.
**CRITICAL: ALWAYS respond in the same language as the user's input/query.**

**PRIMARY OBJECTIVE**: Collect all necessary raw data about a location and store it in session
state for use by downstream analysis and reporting agents.

....

**OUTPUT**:

Your final response should be a raw json containing all the following tags:
- `location_coordinates`: {"lat": float, "lng": float, "formatted_address": string}
- `search_categories`: [list of extracted category strings]
- `nearby_places`: {category: {"count": int, "places": [...]}}
- `existing_charging_stations`: [list of charging station data]
- `traffic_patterns`: {traffic and route analysis data}

**OUTPUT FORMAT EXAMPLE**:
```json
{
  "location_coordinates": {"lat": 48.1351, "lng": 11.5820, "formatted_address": "Munich, Germany"},
  "search_categories": ["gas_station", "shopping_mall", "restaurant"],
  "nearby_places": {"gas_station": {"count": 15, "places": ["Aral", "Shell"]}, "shopping_mall": {"count": 5, "places": ["Riem Arcaden", "Olympia Einkaufszentrum"]}},
  "existing_charging_stations": [{"name": "Supercharger Munich", "location": "XY Street"}],
  "traffic_patterns": [{"time": "8:00", "traffic_volume": "high"}]
}
```

This suggests that gemini-2.0-flash-lite was already aligned closely enough to the teacher that prompt edits had diminishing returns.

Why did this happen?

This can be explained by the interaction of:

A. Strong baseline alignment in medium-sized models

Instruction-following models in the <4B range (including Gemini Lite variants) are often optimized for:

  • schema-aligned generation
  • deterministic JSON/structured outputs
  • short reasoning chains

For constrained, schema-based tasks, the marginal value of a stronger model becomes small.

B. DSPy’s refinement algorithm is stability-preserving (and tunable)

DSPy’s Refine algorithm is intentionally conservative: it prioritizes stability and avoids introducing changes unless there is a consistent error signal. In particular, it avoids:

  • Hallucinating structural changes
  • Overfitting to small or noisy teacher output
  • Aggressively rewriting prompts without clear benefit

That said, Refine is tunable, and several parameters can influence its behavior, including:

  • N: number of module executions per input (higher values increase exploration)
  • reward_fn: the reward function that evaluates outputs
  • threshold: minimum reward needed to accept a prediction

In our experiments, adjusting these parameters, such as increasing N or tightening the evaluation metric, led to earlier convergence rather than more aggressive prompt changes, because the student model already matched the reference outputs with high fidelity.

As a result, Refine correctly inferred that no corrective signal existed, and therefore limited its edits to minimal (ex, adding short examples).

When the gap is already negligible, as in our case, additional tuning does not unlock further gains, and conservative convergence is the correct outcome.

C. Prompt refinement cannot improve what is not broken

Refine added examples because examples are the safest, lowest-risk enhancement.

But since the student model already understood the task, examples did not shift the distribution of outputs.

Understanding the mechanism

This behavior reflects the interplay of generalization vs explicit prompting:

  • When the student already shares the teacher’s inductive biases,
  • And the task is deterministic, structural, and low-entropy,
  • Prompt edits hardly affects the answer.

The model relies on internal generalization, not the exact wording of the prompt.

The experiment, therefore, reveals that teacher–student closeness determines the possible benefit of prompt refinement.

Key insight

When the student model already accurately captures the teacher’s behavior, DSPy’s prompt refinement will have minimal measurable effect, and that is a positive sign of model stability.

The refined and unrefined prompts behave almost identically because the model is already doing the right thing.

Practical implications

Before applying automated prompt optimization (DSPy Refine, PACE, ≥1-shot augmentations, etc.):

  1. Measure the natural output gap between teacher and student

    If the gap is tiny, refinement gains will be limited.

  2. Check whether the task is structured or open-ended

    Structured tasks saturate quickly; creative tasks benefit more from refinement.

  3. Use refinement primarily when correcting errors, not when polishing stable behavior

If your student model is already nearly identical to your teacher, prompt refinement may be optional, not essential.