Skip to content

Feature/policy - #930

Draft
myrazma wants to merge 8 commits into
dev/candidatesfrom
feature/policy
Draft

myrazma wants to merge 8 commits into
dev/candidatesfrom
feature/policy

Conversation

@myrazma

@myrazma myrazma commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator

Policies

A policy is a transformation step that operates on a CandidatesProtocol,
applied before the recommender. It is distinct from a recommender: recommenders select a batch from an already-
materialized, finite candidate set; policies produce or filter the candidate set
before the recommender ever sees it.

The central abstraction is PolicyProtocol. Multiple concrete implementations are
expected; they share a single interface so the search space can call any of them
uniformly.

Examples of policies:

  • Random – randomly sample parameter values from the search space
  • Generator – generate new candidates from infinite parameters
  • Genetic Algorithm – generate new candidates based on previous top performers

Integration Point

The policy is passed at call time of the recommender's .recommend() and gets passed down to SubspaceDiscrete.get_candidates().
A policy is not stored as a property on the campaign, since it may change or be updated each iteration; this also avoids serialization complexity.

    campaign = Campaign(
        searchspace=searchspace, objective=objective, recommender=recommender
    )
    campaign.add_measurements(measurements)
    policy = RandomSamplingPolicy(n=1000)
    recommendations = campaign.recommend(batch_size=16, policy=policy)

The call chain is:

Campaign.recommend(..., policy=self.policy)
  → recommender.recommend(..., policy=self.policy)
    → SubspaceDiscrete.get_candidates(policy=policy)
      → policy(self.candidates, ...)

Responsibilities

A policy must handle two situations:

  1. Infinite candidates (CandidatesProtocol.is_finite is False): The policy must
    materialize the space into a concrete set of Candidates. Without a policy, calling
    to_lazy() on an infinite CandidatesProtocol raises InfiniteSpaceError.
  2. Finite candidates (CandidatesProtocol.is_finite is True): The policy may
    filter, transform, or resample the already-materializable set.

Policy chains

A PolicyChain holds an ordered tuple of policies and applies them in sequence.
A PolicyChain itself implements PolicyProtocol, so from the outside it looks like a single policy — get_candidates() has no knowledge of the internal composition.
Only the first policy in the chain may be a generator policy that operates on an unmaterialized/infinite set; all subsequent policies must operate on a finite set.

Chaining policies enables more advanced workflows, for example:

  1. Generate n sequences from the definition of a sequence parameter using a generative model
  2. Make sure the generated sequences comply with any requirements for the sequence
  3. Subsample n sequences in a user-defined (udf) way, or just randomly for simplicity
img

Relationship to Recommenders

Policy and Recommender are sequential, not interchangeable: policy runs first, recommender runs on the policy's output.
A recommender may be used as a policy, but not vice versa.

Concern Policy Recommender
Materializes infinite space Yes No
Filters/generates candidates Yes No
Scores candidates Optional Yes
Sees measurements Optional Yes
Sees objective Optional Yes
Operates on Continuous Searchspace No Yes

Additions & Roadmap

[X] Implement PolicyProtocol and enable use of a single policy
[X] Implement PolicyChain(PolicyProtocol) to enable policy chain
[ ] Handle the sequential ordering of different kinds of candidates with respect to materialized vs. non-materialized candidates, especially within chaining
[ ] Add a simple RandomPolicy
[ ] Enable a RecommenderPolicy: use model-driven recommenders as policies. A recommender gets the function .to_policy(). Open issue: how to handle the search space at initialization?
[ ] Enable a ParameterPolicy: offer a policy that only works on one parameter (using the parameter selector), with a .to_policy() method to convert it into a Candidates policy.
[ ] Make policies convenient for SequenceParameter: decide how and where to store the values of a SequenceParameter after generation and filtering in policies

Open Todos to finalize PR:
[ ] Critical policy is not passed in all subcalls of get_candidates() yet
[ ] Critical get_candidates() output type change is currently incomplete: Needs to match narwhalified data frames; not applied everywhere; need to check if policy needs to be passed as well.
[ ] Narwhalify! (and remove .to_pandas() tmp workaround)
[ ] Restrict conversion from TableCandidates to ProductCandidates, since that direction loses information (the reverse is fine): in ProductCandidates, we first define the parameters and only at the end call .to_lazy() to get the combinations — we never modify combinations directly. In TableCandidates, we operate directly on the dataframe and can modify candidate combinations.
[ ] Examples
[ ] Tests
[ ] Docs
[ ] Update according to Pyrefly for all code changes

Changes to the Existing Codebase (and Why)

SubspaceDiscrete.get_candidates() should return a CandidatesProtocol instead of a DataFrame.

--- This is not fully integrated yet in all .get_candiate() calls ---
Rationale:
Campaign.recommend() does the following with respect to the discrete search space and its candidates:

candidates_df = self.searchspace.discrete.get_candidates(policy=policy) # The candidates may have changed -> Infinite Sequence Parameter to finite sequence parameter with actual values
mask_todrop = ... # Some constraints & filter actions if requested
searchspace = evolve(
    self.searchspace,
    discrete=evolve(
        self.searchspace.discrete,
        candidates=TableCandidates(
            self.searchspace.discrete.parameters,
            candidates_df.loc[~mask_todrop],
        ),
    ),
)
recommender.recommend(.., searchspace, ..)

Now, what's the issue?
self.searchspace.discrete.get_candidates(policy=policy) can modify both the candidates and the parameters.
If it only returns the dataframe — and not the evolved candidates set with its modified parameters — the discrete search space never learns about the updated parameters and candidates.
What we actually need is the full candidates object returned by .get_candidates(policy), so we can use the parameters from the updated candidates set:

candidates = self.searchspace.discrete.get_candidates(policy=policy) # The candidates may have changed -> Infinite Sequence Parameter to finite sequence parameter with actual values
candidates_df = candidates.to_lazy()
mask_todrop = ... # Some constraints & filter actions if requested
searchspace = evolve(
    self.searchspace,
    discrete=evolve(
        self.searchspace.discrete,
        candidates=TableCandidates(
            candidates.parameters, candidates_df.loc[~mask_todrop] # use from updated canddiates set, not the search space
        ),
    ),
)
recommender.recommend(.., searchspace, ..)

Note how we use candidates.parameters instead of self.searchspace.discrete.parameters to get the updated version.
Changes applied to the candidates during .get_candidates(policy=policy) remain local to that .recommend() call — the Campaign's search space never sees the modified version.
This follows the implementation paradigm "what happens in the recommend call, stays in the recommend call": the down-filtering of the candidates stays in the recommend call and is not propagated back to the campaign's search space.

This need became apparent while working with a parameter and updating its active_values in a policy:
without access to those changes, we risk drawing wrong conclusions and using incorrect parameter definitions in the recommendations.

Design Choices and Thoughts

Arguments of a Policy need to be stored as an attribute

In case a policy has a special argument, we would need to store it as an attribute, to maintain the same call functionality across all custom policies. (e.g. batch size, threshold, objective, measurements)

Issue

If measurements are stored as a property on the policy but need to change or grow in the
middle of a campaign, they will not update automatically.

Policy(Candidates) -> Candidates

The input and output of a policy is a candidates set rather than a raw dataframe. This enables chaining.

PolicyChain is on purpose not called MetaPolicy

MetaRecommender in BayBE selects one recommender from a set based on context.
PolicyChain applies all policies in sequence — a fundamentally different operation.
The name MetaPolicy would imply context-dependent selection, which is a separate
concept (and may be added later independently).

_RecommenderPolicy

A policy, in essence, takes a candidates set in and returns a reduced/modified candidates set.
The same applies to a recommender, which reduces the a set of candidates (going from searchspace to recommendations).
We therefore also want to enable using a recommender as a policy — useful for lightweight models and proxy targets that perform a first pre-selection.

An adapter would allow a recommender to act as a policy: call .to_policy(..) with the same signature as .recommend(), which returns a PolicyProtocol.

Issue

A recommender can act as a pre-selection step in a policy chain, but RecommenderProtocol.recommend() takes batch_size, searchspace, objective, measurements, and pending_experiments at call time. PolicyProtocol.__call__ takes only candidates. The adapter must bridge this gap.

Search space
The recommender operates on a search space, while the policy operates on candidates.

Option 1: Infer the search space from the candidates. Limitations: (1) only the discrete
search space is available, and (2) batch constraints are lost, so calling the same
recommender as a policy behaves differently than calling it directly as a recommender —
confusing for the user and a potential source of silent errors.

Option 2: Pass the search space to .to_policy() at construction time and store it as an
attribute on the policy. Issue: the policy operates on incoming candidates via
__call__(); if we store the search space as an attribute, it can drift significantly
from the candidates actually passed in once chaining is involved.

ParameterPolicy:

We want to offer a policy that operates on a single parameter (via the parameter selector) rather than the full candidates set, which requires a way to go from parameter-level operations to a candidates-level policy.

Limitations and Open Points

Refactoring needed before policies can be merged

  • Lazification of TableCandidates: it should accept a LazyFrame so that the policy chain can stay directly in the lazy regime and never move to eager evaluation if not needed. validate_parameter_input also has to work on lazy frames, otherwise we cannot stay in the lazy data regime.
  • More for Sequence parameter:comp_rep_columns needs refactoring to be dynamic - ._make_input_scaler() makes use of comp_rep_columns, which we don't have at hand for the embeddings before runtime. Refactor to be dynamic.

Thoughts on SequenceParameter

Notes on SequenceParameter gathered while testing it with policies:

  • How is the use intended? IMO: SequenceParameter provides the definition and encoding; in theory, values will be infinite, but we can select suitable values by setting active_values and use them as candidates. My suggestion: values - all potential candidates (can be infinite), active_values - recommendable values.
    • Why not use CustomParameter instead? -> Moving to CustomParameter would require encoding the candidates immediately. But in a policy chain, we may want to do further operations on the sequence first, like random subsampling to heavily reduce the set, and only encode at the end. If the encoder for embeddings lived on CustomParameter, what would SequenceParameter even be for?
    • Why not just set the values directly? -> Possible in practice, but conceptually incorrect, since the values set is meant to stay infinite — we still need a way to express and store the possible values for that parameter.
  • comp_rep_columns: Bigger refactoring is coming to make the column names dynamic.
  • Scaling is currently done with all values; this is not feasible for SequenceParameter, as values is always infinite in the current implementation --> Needs refactoring as well.

To Be Discussed

  • How should the flag is_finite be treated for SequenceParameter? The underlying Parameter can be infinite, but by setting active_values it can be made materializable, i.e. infinite != materializable. How should this be handled when working with SequenceParameter and Policies?

get_candiates returned a pd.DataFrame before, when returning a Candidates Protocol,
we can use more information to correctly evolve our search space in recommend().
In addition, we can use the lazificaiton of the specific candidates set.

# Conflicts:
#	baybe/recommenders/pure/bayesian/botorch/discrete.py

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant