Abstract
Large Language Models (LLMs) have been used for novel discovery of algorithms, theorems, drugs, and other tasks through the use of harnesses that prompt an LLM to iteratively optimize an objective. In this work, we study the relationship between the population of previous iterates and eventual discovery success. We generalize past work on harness design to develop a suite of 12 harnesses called 'Modular' and characterize their performance across 5 diverse discovery tasks, finding that discovery success is brittle and sensitive to harness design. We uncover mode collapse, characterized by a dramatic drop in the diversity of iterates, as a common failure mode. We find that popular state-of-the-art harnesses and diversity-inducing harness interventions, which aim to prolong this collapse, yield inconsistent gains. Our results instead uncover that the performance of early discoveries is predictive of eventual success. We therefore propose a universally applicable intervention that performs an initial stage of parallel exploration in order to initialize subsequent iterative optimization. Our method provides consistent gains across many harnesses and target applications, confirming the importance of initialization in LLM-driven discovery.
What is LLM-Driven Discovery?
LLMs can be prompted to 1) generate and 2) iteratively refine "discoveries" (e.g., algorithms, molecules, theorems).
For example, an LLM can be prompted to: "Generate a fast sorting algorithm."
The initially generated algorithm can be timed and tested for correctness, and then this feedback can be provided to the LLM to aid in the generation of the next variant of the sorting algorithm.
The intuition is that providing an LLM feedback on past discovery iterates, can enable better future iterates.
Over the course of a discovery trajectory, an LLM can generate many different discovery iterates. A discovery harness can be used to manage this population of past dicoveries.
Modular: A Suite of Simple Discovery Harnesses
Generally, a harness samples parent discoveries from an active population of past discoveries. An LLM is then prompted to mutate these parents to generate an improved discovery, with respect to the objective. New discoveries undergo task-specific evaluation and are added to the active population if they meet certain criteria. The active population is periodically pruned when it reaches capacity.
# A general discovery harness
argument descriptions:
task_objective: natural language description of discovery task (e.g., "Generate a sorting algorithm")
evaluate_discovery(): a procudure to evaluate an LLM generated discovery
population = [] # active population of past dicoveries
while tokens_used < token_budget:
parent_discoveries = select_parents(population)
new_discovery = LLM_Mutator(parent_dicoveries, task_objective)
score, meta_data = evaluate_discovery(new_discovery)
population = update_population(new_discovery, score, meta_data)
best_discovery = highest_score(population)
We develop a set of harnesses called 'Modular' by generalizing the OPRO discovery framework proposed by Yang et al. (2024). We minimally modify the OPRO harness to expose two axes of freedom: 1) the LLM mutation strategy and 2) the parent sampling strategy to enable a broader range of harness behavior. We consider four mutation strategies and three sampling strategies, resulting in a suite of 12 unique harnesses (detailed in the paper).
Parametric vs. In-Context Knowledge
Parametric knowledge is the knowledge that an LLM stores in its weights; this knowledge is typically acquired during training via gradient-based updates to the LLM's weights (i.e., parameters).
This parametric knowledge can be utilized to accomplish tasks via prompting.
A prompt to a LLM contains in-context knowledge, i.e., knowledge external to the LLM's weights that will have a downstream affect on what the LLM generates.
Here, we will study the role of both parametric and in-context knowledge for LLM-driven discovery.
First, we will use a strightforward method: an initial seed iterate is mutated in parallel by an LLM prompted with the task’s
discovery objective until the token budget is exhausted. Here, the LLM’s context never changes; we call this method 'Parallel'. We compare 'parallel' to the sequential Modular harnesses (described above).
Sequential harnesses employ an LLM to mutate discoveries, but instead iteratively update the LLM’s context with information from recent discoveries.
Since parallel discovery never updates the LLM’s context, its performance is dependent on
the LLM’s parametric knowledge base. In parallel discovery, the ceiling for discovery performance is
not known a priori, as reliably measuring the LLM’s parametric knowledge for discovery tasks is challenging. For example, in the above figure we observe that for TSP, parallel discovery typically underperforms
sequential Modular harnesses, while for Circle Packing the trend is reversed. On the other hand, sequential harnesses enable LLMs to build on the in-context knowledge from past discoveries, rather
than relying solely on parametric knowledge. For each task, we observed that multiple Modular harnesses outperformed parallel discovery, but that the specific outperforming Modular variants vary. This
suggests that sequential refinement of a discovery (with in-context knowledge) can outperform parallel optimization.
Can and Should we Prevent Mode Collapse?
Sequential LLM disocovery harnesses are prone to mode collapse: the diversity of new iterates decreases over the course of a sequential trajectory. We hypothesize that this is due to the in-context prior, which is present and reinforced throughout sequential discovery, but absent in parallel discovery. We find that the parallel schemes generally have more diverse populations of discoveries than sequential harnesses.
We wondered if preventing mode collapse could enable better discoveries. To test this, we compared our simple Modular harnesses to harnesses which prioriotize diversity of iterates. We did this by:
- Comparing Modular harnesses to three state-of-the-art harnesses that treat population diversity as a key consideration in harness design.
- Applying a set of six interventions to Modular harnesses that aim to maintain diversity.
We observe that while high-performing trajectories can be marginally more diverse early on, they too can experience mode collapse over time. This suggests a more nuanced conclusion: in a trajectory with poor initial candidates, mode collapse around those candidates can prevent better ones from ever being discovered, so intervening on the population can help. In contrast, an initially high-performing trajectory may actually be hurt by unnecessary exploration and benefit from more exploitation. Distinguishing between these two scenarios is challenging, and exploration does not guarantee recovering from poor initial discoveries.
Initialization Raises the Discovery Floor
We find that the quality of the initially discovered population can be predictive of
downstream discovery quality in a trajectory. We analyze 75 different discovery harness variants (from the prevous section) below.
We visualize the relationship between the initial score
of the first 15 discoveries in a discovery trajectory vs. the maximum downstream score achieved by that trajectory. While
no specific harness consistently performed well, the highest-performing harnesses for each (task, LLM) pairing tended to be those that produced a high-performing
initial population.
Since the eventual success of a sequential trajectory can be correlated with the performance of the first
few discovered candidates, we posit that it is beneficial to initialize a high-performing population
of candidates that can better condition downstream sequential optimization. We propose a simple and extensible initialization method
that can be adopted with any sequential discovery harness: first generate a pool of discoveries using
parallel discovery with a fraction of the discovery budget (r); then curate a performant population pinit of size m by greedily selecting candidates
from that parallel pool. The intialized pool can be sequentially refined witha discovery harness with the remaining discovery budget.
We plot the average maximal score achieved across all 12 Modular
variants across different initialization budgets. In all settings, we find that reframing the purpose of parallel
discovery from a technique that aims to push the discovery ceiling higher (Init. Budget = 100%) to an initialization
method that raises the discovery floor (0% < Init. Budget < 100%) can be beneficial. Initializing a
serial trajectory with a high-performing population consistently outperforms both purely serial (Init. Budget = 0%)
and purely parallel (Init. Budget = 100%) discovery, given enough additional serial compute. We find that parallel discovery typically converges in discovery quality, and
therefore propose using convergence as a simple indicator of the ideal initialization budget (see details in paper).
Finally, we validate that explicit population initialization outperforms the implicit initialization achieved by SOTA harnesses above.
Conclusion
In this work, we study the relationship between how discovery harnesses are initialized and how this affects downstream discovery quality. Through the development and characterization of a suite of discovery harnesses, we validate that sequential refinement of an initial seed discovery is beneficial, but sensitive to harness design. Subsequently, we find that harness design choices that aim to enable more diverse discoveries, from an initial seed discovery, do not predictably achieve more performant discoveries. However, we notice, across a wide range of (task, LLM) settings, that the best performing discovery trajectories can be predicted by their initial discovery quality. Building on this, we propose a simple, task-, LLM-, and harness-agnostic method to curate high-performing initial discovery populations. By warm-starting discovery harnesses with these initial populations, we consistently improve discovery harnesses’ performance.
BibTeX
@article{sakarvadia2026initialization,
title={Initialization Improves LLM-Driven Discovery},
author={Mansi Sakarvadia and Marco Ciccone and Colin Raffel},
year={2026},
url={https://arxiv.org/abs/2610.00707}
}