How AI Agents Choose What to Learn Next: Expected Information Gain Explained

Utpal Kumar   8 minute read      

Conceptual illustration of an AI research agent choosing a sensor observation while considering three possible explanations for a building’s changed vibration, represented by a thermometer, weights, and a structural beam.

A monitoring dashboard shows that a building’s natural frequency has fallen. Several explanations are plausible: a temperature-related change, extra mass inside the building, or a persistent loss of stiffness. Another hour of vibration data might confirm the shift. It might still leave us unsure what caused it.

An AI agent helping with the investigation has a choice. It could retrieve temperature records, inspect loading information, or compare another vibration mode. Each action returns something, but some results would be more useful than others.

This is the question behind expected information gain, usually shortened to EIG: before collecting more evidence, which observation is likely to reduce our uncertainty the most?

In my article on agentic harnesses, I described the software that connects a model to tools and feedback. Here, I want to look at one way of choosing the next step within that loop. EIG is a principle from Bayesian experimental design that we can use to think more carefully about what an agent should investigate.

Start with the uncertainty you want to resolve

The building example contains two different questions. Has its frequency changed? And what explains the change?

A longer recording may answer the first question by giving a more precise frequency estimate. Answering the second requires observations that help distinguish the competing explanations.

Temperature matters because it can affect material properties and the restraint supplied by supports and joints. Added mass can also lower a natural frequency, as can reduced stiffness. The Z24 Bridge study by Peeters and De Roeck illustrates why environmental variation must be considered when interpreting frequency changes.

Suppose all three explanations predict roughly the same drop in the first vibration mode. Repeating that observation may tell us little about which explanation is right. A temperature history or a measurement of another mode could be more informative if the explanations predict different results for it.

That last condition matters. A measurement becomes useful through its relationship to the question. A large signal, a precise sensor, or a long record does not automatically provide a strong distinction between causes.

For this article, the target is uncertainty about the explanation. In another task, it might be an earthquake’s location, a material parameter, or the reason a software test is failing. Define that target before ranking possible observations.

Expected means considering the possible outcomes

We choose a measurement before seeing its result. EIG asks how much uncertainty we expect to remove, averaging over the results our model considers possible.

The uncertainty measure is Shannon entropy. If pᵢ is the probability assigned to explanation i, the formula is:

H = −∑ᵢ pᵢ log₂(pᵢ)

The sum runs over all the explanations. A probability of zero contributes zero. Using base-two logarithms gives entropy its unit: bits.

For n equally likely explanations, every pᵢ equals 1/n. Substituting into the formula gives:

H = −n × (1/n) × log₂(1/n) = log₂ n

The base-two logarithm asks what power of 2 gives n. The numbers follow directly:

  • One certain explanation: H = log₂(1) = 0 bits.
  • Two equally likely explanations: H = log₂(2) = 1 bit.
  • Three equally likely explanations: H = log₂(3) ≈ 1.585 bits.

If the explanations are not equally likely, use the weighted sum in the first formula.

To calculate EIG for a candidate measurement:

  1. List its possible results and predict how likely each is, using the current hypothesis probabilities and what each hypothesis predicts.
  2. For each result, update the hypothesis probabilities as though you had observed it. Calculate the entropy of those updated probabilities.
  3. Multiply each result’s probability by the entropy it would leave. Add these products to get the average remaining uncertainty.
  4. Subtract that average from the entropy before the measurement.

Expected information gain = uncertainty now − average uncertainty after the measurement.

For a preview of the example below, suppose a check has a one-third chance of leaving zero bits and a two-thirds chance of leaving one bit. Its average remaining uncertainty is (1/3 × 0) + (2/3 × 1) = 0.667 bits. Starting from 1.585 bits, its EIG is 1.585 − 0.667 = 0.918 bits.

These probabilities come from the assumed model. They describe possible future results, not the agent’s confidence in its prose. This expected entropy reduction is also called mutual information; the Bayesian Active Learning by Disagreement paper develops the connection for informative query selection.

A small example makes the choice visible

Assume exactly one of three explanations is true, each with probability one-third:

  • A reversible temperature-related effect
  • Added mass
  • A persistent stiffness change

These probabilities are invented for teaching. Real causes can coexist, and real measurements are noisy.

Check A repeats the frequency measurement under the same conditions. All three explanations predict the same low frequency with certainty in this toy model. Seeing it again leaves their probabilities unchanged, so EIG is zero.

Check B measures the frequency after the building’s temperature returns to its reference state. Suppose earlier observations established a reversible temperature–frequency relationship. We compare the new frequency with the baseline measured at that same reference state. Does the frequency recover to baseline, or remain low?

Assume the building has fully returned to that thermal state, any added mass remains in place, any persistent stiffness change remains unchanged, and the measurement is error-free. This idealized model predicts:

  • Temperature explanation: frequency recovers with probability 1; it remains low with probability 0.
  • Added-mass explanation: frequency recovers with probability 0; it remains low with probability 1.
  • Persistent-stiffness explanation: frequency recovers with probability 0; it remains low with probability 1.

Only the temperature explanation predicts recovery. Its one-third probability gives recovery a one-third chance. The other two explanations both predict a low frequency, giving that outcome a combined two-thirds chance.

If frequency recovers, only the temperature explanation fits: its updated probability is 1 and the others are 0. The remaining entropy is zero bits.

If frequency stays low, the temperature explanation is ruled out within this model. The two surviving explanations started equally likely and predict this result equally well, so each now has probability one-half. The remaining entropy is one bit.

Weight the remaining uncertainties by one-third and two-thirds. Their average is 0.667 bits. Subtract it from the starting 1.585 bits to obtain 0.918 bits of expected information gain.

The more likely outcome still leaves two explanations. Check B is useful because either result narrows the investigation. In a real building, recovery would be evidence to weigh alongside other observations, not proof of a single cause or of safety. Thermal lag, changing loads, and measurement errors would require less decisive probabilities.

Illustrative idealized comparison after a building’s vibration frequency drops. Three mutually exclusive explanations start equally likely: temperature-related effect T, added mass M, and persistent stiffness change S, each with probability one third. Check A repeats the frequency measurement under unchanged conditions and leaves all three equally likely. Check B measures again after the building returns to its reference thermal state, assuming unchanged added mass and persistent stiffness change and no measurement error. Recovery has probability one third and leaves only T in this model. A frequency that stays low has probability two thirds and leaves M and S equally likely. Expected information gain is about 0.92 bits.
Figure 1. An illustrative, idealized comparison. Check B observes whether frequency recovers when the building returns to its reference thermal state. Recovery leaves the temperature explanation; a low frequency leaves added mass and persistent stiffness change equally likely. EIG is about 0.92 bits. Real checks require noise and overlapping causes to be modeled.

Applying the idea to an agent’s next tool call

An agent investigating a failed data-processing job faces a similar problem. Perhaps the input file is missing, its schema has changed, or the code expects the wrong units.

Reading an unrelated source file may return plenty of text. Checking whether the input exists separates one explanation from the others. Inspecting the header may distinguish a schema problem. Comparing units against the data specification may resolve a different uncertainty.

An EIG-guided loop would work like this:

  1. State the unresolved question and the explanations currently supported by evidence.
  2. List the permitted tool calls that could help answer it.
  3. Predict plausible results for each call under those explanations.
  4. Estimate how much each result would change the uncertainty, and average over the results.
  5. Run a worthwhile call, inspect the actual output, and update before choosing again.

The ReAct paper demonstrates reasoning and acting interleaved with observations. It does not establish that those agents calculate EIG. The connection here is a design possibility: information gain supplies a criterion for choosing an information-gathering action within such a loop.

A numerical implementation needs more than a language model saying it is “90% confident.” It needs explicit hypotheses or parameters, defensible probabilities, and a model of what each tool might return, including failures. Those probabilities should be checked against representative cases.

Without that machinery, an agent can still ask, “Which result would distinguish the remaining explanations?” That is a useful heuristic, but it should not be presented as a calibrated EIG calculation.

An expected-information-gain planning loop with six steps. Start with current evidence and beliefs. List permitted candidate tool calls. Predict each call’s possible outcomes. Estimate information gain by averaging the uncertainty reduction over those outcomes. Choose an informative call while accounting for cost, risk, and constraints. Observe the real result and update beliefs, then return to the start. Expected information gain is a prediction made before acting, not a guarantee of what one observation will reveal.
Figure 2. A conceptual information-seeking loop. Current evidence supports predictions about possible tool results. The system chooses an allowed action, observes what actually happens, and updates. Costs, permissions, and stopping conditions constrain the loop throughout.

Knowing when more information is worth getting

EIG depends on the model used to calculate it. If our building model omits a sensor fault, the system may become increasingly certain about the wrong explanation. Research on experimental design under model misspecification addresses this gap between an assumed model and the process producing the observations. In practice, inspect unexpected results and revisit the explanation set when the evidence fits it poorly.

Calculation also has a cost. Complex problems may require many simulated outcomes and difficult probability estimates for each candidate action. Modern Bayesian Experimental Design reviews these computational challenges and the approximations used to address them. An exact score is rarely something we get for free.

Finally, information is only one part of a decision. A slow, expensive measurement may reduce uncertainty more than a quick check while arriving too late to help. Learning a minor parameter very precisely may contribute little to the action we need to take. Where a decision and its consequences are known, expected improvement in that decision is often the more direct objective.

For an agent, this means considering latency, tool cost, reliability, and the task’s completion criteria alongside information gain. A high score cannot authorize a tool call outside its permissions. For structural monitoring, collecting more data must never delay an inspection or protective action that is already warranted; a frequency-based example cannot establish a building’s safety.

The practical habit is simple. Before requesting another dataset, running another tool, or installing another sensor, name the uncertainty that matters. Write down what the remaining explanations predict. Then look for an observation whose possible outcomes would help separate them.

That is the useful question EIG makes precise: what could we learn from the next result, and how likely are we to learn it?

Disclaimer of liability

The information provided by the Earth Inversion is made available for educational purposes only.

Whilst we endeavor to keep the information up-to-date and correct. Earth Inversion makes no representations or warranties of any kind, express or implied about the completeness, accuracy, reliability, suitability or availability with respect to the website or the information, products, services or related graphics content on the website for any purpose.

UNDER NO CIRCUMSTANCE SHALL WE HAVE ANY LIABILITY TO YOU FOR ANY LOSS OR DAMAGE OF ANY KIND INCURRED AS A RESULT OF THE USE OF THE SITE OR RELIANCE ON ANY INFORMATION PROVIDED ON THE SITE. ANY RELIANCE YOU PLACED ON SUCH MATERIAL IS THEREFORE STRICTLY AT YOUR OWN RISK.


Leave a comment