Initialization Improves LLM-Driven Discovery

1University of Chicago, 2Vector Institute, 3University of Toronto
Initializing discovery harnesses improves LLM-driven discovery.

Initialization raises the floor for discovery under a fixed total token budget. We find that initializing (the Modular) discovery harnesses with a high-performing initial population is beneficial for downstream discovery success across harnesses, tasks, and LLMs.

Abstract

Large Language Models (LLMs) have been used for novel discovery of algorithms, theorems, drugs, and other tasks through the use of harnesses that prompt an LLM to iteratively optimize an objective. In this work, we study the relationship between the population of previous iterates and eventual discovery success. We generalize past work on harness design to develop a suite of 12 harnesses called 'Modular' and characterize their performance across 5 diverse discovery tasks, finding that discovery success is brittle and sensitive to harness design. We uncover mode collapse, characterized by a dramatic drop in the diversity of iterates, as a common failure mode. We find that popular state-of-the-art harnesses and diversity-inducing harness interventions, which aim to prolong this collapse, yield inconsistent gains. Our results instead uncover that the performance of early discoveries is predictive of eventual success. We therefore propose a universally applicable intervention that performs an initial stage of parallel exploration in order to initialize subsequent iterative optimization. Our method provides consistent gains across many harnesses and target applications, confirming the importance of initialization in LLM-driven discovery.

What is LLM-Driven Discovery?

LLMs can be prompted to 1) generate and 2) iteratively refine "discoveries" (e.g., algorithms, molecules, theorems).

For example, an LLM can be prompted to: "Generate a fast sorting algorithm." The initially generated algorithm can be timed and tested for correctness, and then this feedback can be provided to the LLM to aid in the generation of the next variant of the sorting algorithm. The intuition is that providing an LLM feedback on past discovery iterates, can enable better future iterates.

Over the course of a discovery trajectory, an LLM can generate many different discovery iterates. A discovery harness can be used to manage this population of past dicoveries.

Modular: A Suite of Simple Discovery Harnesses

Generally, a harness samples parent discoveries from an active population of past discoveries. An LLM is then prompted to mutate these parents to generate an improved discovery, with respect to the objective. New discoveries undergo task-specific evaluation and are added to the active population if they meet certain criteria. The active population is periodically pruned when it reaches capacity.


# A general discovery harness
					
argument descriptions:
	task_objective: natural language description of discovery task (e.g., "Generate a sorting algorithm")
	evaluate_discovery(): a procudure to evaluate an LLM generated discovery

population = [] # active population of past dicoveries	
while tokens_used < token_budget:
	parent_discoveries = select_parents(population)
	new_discovery = LLM_Mutator(parent_dicoveries, task_objective)
	score, meta_data = evaluate_discovery(new_discovery)
	population = update_population(new_discovery, score, meta_data)

best_discovery = highest_score(population)
				

We develop a set of harnesses called 'Modular' by generalizing the OPRO discovery framework proposed by Yang et al. (2024). We minimally modify the OPRO harness to expose two axes of freedom: 1) the LLM mutation strategy and 2) the parent sampling strategy to enable a broader range of harness behavior. We consider four mutation strategies and three sampling strategies, resulting in a suite of 12 unique harnesses (detailed in the paper).

Parametric vs. In-Context Knowledge

Parametric knowledge is the knowledge that an LLM stores in its weights; this knowledge is typically acquired during training via gradient-based updates to the LLM's weights (i.e., parameters). This parametric knowledge can be utilized to accomplish tasks via prompting. A prompt to a LLM contains in-context knowledge, i.e., knowledge external to the LLM's weights that will have a downstream affect on what the LLM generates.

Here, we will study the role of both parametric and in-context knowledge for LLM-driven discovery. First, we will use a strightforward method: an initial seed iterate is mutated in parallel by an LLM prompted with the task’s discovery objective until the token budget is exhausted. Here, the LLM’s context never changes; we call this method 'Parallel'. We compare 'parallel' to the sequential Modular harnesses (described above). Sequential harnesses employ an LLM to mutate discoveries, but instead iteratively update the LLM’s context with information from recent discoveries.
Sequential discovery can outperform parallel discovery.
Since parallel discovery never updates the LLM’s context, its performance is dependent on the LLM’s parametric knowledge base. In parallel discovery, the ceiling for discovery performance is not known a priori, as reliably measuring the LLM’s parametric knowledge for discovery tasks is challenging. For example, in the above figure we observe that for TSP, parallel discovery typically underperforms sequential Modular harnesses, while for Circle Packing the trend is reversed. On the other hand, sequential harnesses enable LLMs to build on the in-context knowledge from past discoveries, rather than relying solely on parametric knowledge. For each task, we observed that multiple Modular harnesses outperformed parallel discovery, but that the specific outperforming Modular variants vary. This suggests that sequential refinement of a discovery (with in-context knowledge) can outperform parallel optimization.

Can and Should we Prevent Mode Collapse?

Sequential LLM disocovery harnesses are prone to mode collapse: the diversity of new iterates decreases over the course of a sequential trajectory. We hypothesize that this is due to the in-context prior, which is present and reinforced throughout sequential discovery, but absent in parallel discovery. We find that the parallel schemes generally have more diverse populations of discoveries than sequential harnesses.

Sequential discovery harnesses suffer from mode collapse.

We wondered if preventing mode collapse could enable better discoveries. To test this, we compared our simple Modular harnesses to harnesses which prioriotize diversity of iterates. We did this by:

  1. Comparing Modular harnesses to three state-of-the-art harnesses that treat population diversity as a key consideration in harness design.
  2. Applying a set of six interventions to Modular harnesses that aim to maintain diversity.
Our results from state-of-the-art harnesses and diversity interventions show that when harnesses start from the same seed, those that explore more (i.e., maintain more diverse populations) do not reliably produce better-performing discoveries. For example, in we visually stratify all 72 discovery trajectories for the CloudCast task into tertiles by their best score.
High scoring trajectories have higher initial diversity but eventually experience mode collapse.
We observe that while high-performing trajectories can be marginally more diverse early on, they too can experience mode collapse over time. This suggests a more nuanced conclusion: in a trajectory with poor initial candidates, mode collapse around those candidates can prevent better ones from ever being discovered, so intervening on the population can help. In contrast, an initially high-performing trajectory may actually be hurt by unnecessary exploration and benefit from more exploitation. Distinguishing between these two scenarios is challenging, and exploration does not guarantee recovering from poor initial discoveries.

Initialization Raises the Discovery Floor

We find that the quality of the initially discovered population can be predictive of downstream discovery quality in a trajectory. We analyze 75 different discovery harness variants (from the prevous section) below.
Initial discovery quality can predict downstream quality.
We visualize the relationship between the initial score of the first 15 discoveries in a discovery trajectory vs. the maximum downstream score achieved by that trajectory. While no specific harness consistently performed well, the highest-performing harnesses for each (task, LLM) pairing tended to be those that produced a high-performing initial population.

Since the eventual success of a sequential trajectory can be correlated with the performance of the first few discovered candidates, we posit that it is beneficial to initialize a high-performing population of candidates that can better condition downstream sequential optimization. We propose a simple and extensible initialization method that can be adopted with any sequential discovery harness: first generate a pool of discoveries using parallel discovery with a fraction of the discovery budget (r); then curate a performant population pinit of size m by greedily selecting candidates from that parallel pool. The intialized pool can be sequentially refined witha discovery harness with the remaining discovery budget.

Initialized discovery can outperform purely sequential and parallel discovery.
We plot the average maximal score achieved across all 12 Modular variants across different initialization budgets. In all settings, we find that reframing the purpose of parallel discovery from a technique that aims to push the discovery ceiling higher (Init. Budget = 100%) to an initialization method that raises the discovery floor (0% < Init. Budget < 100%) can be beneficial. Initializing a serial trajectory with a high-performing population consistently outperforms both purely serial (Init. Budget = 0%) and purely parallel (Init. Budget = 100%) discovery, given enough additional serial compute. We find that parallel discovery typically converges in discovery quality, and therefore propose using convergence as a simple indicator of the ideal initialization budget (see details in paper).

Initializing state-of-the-art harnesses improves discoveries.
Finally, we validate that explicit population initialization outperforms the implicit initialization achieved by SOTA harnesses above.

Conclusion

In this work, we study the relationship between how discovery harnesses are initialized and how this affects downstream discovery quality. Through the development and characterization of a suite of discovery harnesses, we validate that sequential refinement of an initial seed discovery is beneficial, but sensitive to harness design. Subsequently, we find that harness design choices that aim to enable more diverse discoveries, from an initial seed discovery, do not predictably achieve more performant discoveries. However, we notice, across a wide range of (task, LLM) settings, that the best performing discovery trajectories can be predicted by their initial discovery quality. Building on this, we propose a simple, task-, LLM-, and harness-agnostic method to curate high-performing initial discovery populations. By warm-starting discovery harnesses with these initial populations, we consistently improve discovery harnesses’ performance.

BibTeX

@article{sakarvadia2026initialization,
  title={Initialization Improves LLM-Driven Discovery},
  author={Mansi Sakarvadia and Marco Ciccone and Colin Raffel},
  year={2026},
  url={https://arxiv.org/abs/2610.00707}
}