Repository navigation
Feature/policy - #930
Draft
myrazma wants to merge 8 commits into
Draft
Feature/policy#930myrazma wants to merge 8 commits into
myrazma wants to merge 8 commits into
Conversation
get_candiates returned a pd.DataFrame before, when returning a Candidates Protocol, we can use more information to correctly evolve our search space in recommend(). In addition, we can use the lazificaiton of the specific candidates set. # Conflicts: # baybe/recommenders/pure/bayesian/botorch/discrete.py
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Policies
A policy is a transformation step that operates on a
CandidatesProtocol,applied before the recommender. It is distinct from a recommender: recommenders select a batch from an already-
materialized, finite candidate set; policies produce or filter the candidate set
before the recommender ever sees it.
The central abstraction is
PolicyProtocol. Multiple concrete implementations areexpected; they share a single interface so the search space can call any of them
uniformly.
Examples of policies:
Integration Point
The policy is passed at call time of the recommender's
.recommend()and gets passed down toSubspaceDiscrete.get_candidates().A policy is not stored as a property on the campaign, since it may change or be updated each iteration; this also avoids serialization complexity.
The call chain is:
Responsibilities
A policy must handle two situations:
CandidatesProtocol.is_finite is False): The policy mustmaterialize the space into a concrete set of
Candidates. Without a policy, callingto_lazy()on an infiniteCandidatesProtocolraisesInfiniteSpaceError.CandidatesProtocol.is_finite is True): The policy mayfilter, transform, or resample the already-materializable set.
Policy chains
A
PolicyChainholds an ordered tuple of policies and applies them in sequence.A
PolicyChainitself implementsPolicyProtocol, so from the outside it looks like a single policy —get_candidates()has no knowledge of the internal composition.Only the first policy in the chain may be a generator policy that operates on an unmaterialized/infinite set; all subsequent policies must operate on a finite set.
Chaining policies enables more advanced workflows, for example:
Relationship to Recommenders
Policy and Recommender are sequential, not interchangeable: policy runs first, recommender runs on the policy's output.
A recommender may be used as a policy, but not vice versa.
Additions & Roadmap
[X] Implement
PolicyProtocoland enable use of a single policy[X] Implement
PolicyChain(PolicyProtocol)to enable policy chain[ ] Handle the sequential ordering of different kinds of candidates with respect to materialized vs. non-materialized candidates, especially within chaining
[ ] Add a simple RandomPolicy
[ ] Enable a RecommenderPolicy: use model-driven recommenders as policies. A recommender gets the function
.to_policy(). Open issue: how to handle the search space at initialization?[ ] Enable a ParameterPolicy: offer a policy that only works on one parameter (using the parameter selector), with a
.to_policy()method to convert it into a Candidates policy.[ ] Make policies convenient for SequenceParameter: decide how and where to store the values of a SequenceParameter after generation and filtering in policies
Open Todos to finalize PR:
[ ] Critical policy is not passed in all subcalls of get_candidates() yet
[ ] Critical get_candidates() output type change is currently incomplete: Needs to match narwhalified data frames; not applied everywhere; need to check if policy needs to be passed as well.
[ ] Narwhalify! (and remove .to_pandas() tmp workaround)
[ ] Restrict conversion from
TableCandidatestoProductCandidates, since that direction loses information (the reverse is fine): inProductCandidates, we first define the parameters and only at the end call.to_lazy()to get the combinations — we never modify combinations directly. InTableCandidates, we operate directly on the dataframe and can modify candidate combinations.[ ] Examples
[ ] Tests
[ ] Docs
[ ] Update according to Pyrefly for all code changes
Changes to the Existing Codebase (and Why)
SubspaceDiscrete.get_candidates()should return aCandidatesProtocolinstead of a DataFrame.--- This is not fully integrated yet in all
.get_candiate()calls ---Rationale:
Campaign.recommend()does the following with respect to the discrete search space and its candidates:Now, what's the issue?
self.searchspace.discrete.get_candidates(policy=policy)can modify both the candidates and the parameters.If it only returns the dataframe — and not the evolved candidates set with its modified parameters — the discrete search space never learns about the updated parameters and candidates.
What we actually need is the full candidates object returned by
.get_candidates(policy), so we can use the parameters from the updated candidates set:Note how we use
candidates.parametersinstead ofself.searchspace.discrete.parametersto get the updated version.Changes applied to the candidates during
.get_candidates(policy=policy)remain local to that.recommend()call — the Campaign's search space never sees the modified version.This follows the implementation paradigm "what happens in the recommend call, stays in the recommend call": the down-filtering of the candidates stays in the recommend call and is not propagated back to the campaign's search space.
This need became apparent while working with a parameter and updating its
active_valuesin a policy:without access to those changes, we risk drawing wrong conclusions and using incorrect parameter definitions in the recommendations.
Design Choices and Thoughts
Arguments of a Policy need to be stored as an attribute
In case a policy has a special argument, we would need to store it as an attribute, to maintain the same call functionality across all custom policies. (e.g. batch size, threshold, objective, measurements)
Issue
If measurements are stored as a property on the policy but need to change or grow in the
middle of a campaign, they will not update automatically.
Policy(Candidates) -> Candidates
The input and output of a policy is a candidates set rather than a raw dataframe. This enables chaining.
PolicyChainis on purpose not calledMetaPolicyMetaRecommenderin BayBE selects one recommender from a set based on context.PolicyChainapplies all policies in sequence — a fundamentally different operation.The name
MetaPolicywould imply context-dependent selection, which is a separateconcept (and may be added later independently).
_RecommenderPolicyA policy, in essence, takes a candidates set in and returns a reduced/modified candidates set.
The same applies to a recommender, which reduces the a set of candidates (going from searchspace to recommendations).
We therefore also want to enable using a recommender as a policy — useful for lightweight models and proxy targets that perform a first pre-selection.
An adapter would allow a recommender to act as a policy: call
.to_policy(..)with the same signature as.recommend(), which returns aPolicyProtocol.Issue
A recommender can act as a pre-selection step in a policy chain, but
RecommenderProtocol.recommend()takesbatch_size,searchspace,objective,measurements, andpending_experimentsat call time.PolicyProtocol.__call__takes onlycandidates. The adapter must bridge this gap.Search space
The recommender operates on a search space, while the policy operates on candidates.
Option 1: Infer the search space from the candidates. Limitations: (1) only the discrete
search space is available, and (2) batch constraints are lost, so calling the same
recommender as a policy behaves differently than calling it directly as a recommender —
confusing for the user and a potential source of silent errors.
Option 2: Pass the search space to
.to_policy()at construction time and store it as anattribute on the policy. Issue: the policy operates on incoming candidates via
__call__(); if we store the search space as an attribute, it can drift significantlyfrom the candidates actually passed in once chaining is involved.
ParameterPolicy:We want to offer a policy that operates on a single parameter (via the parameter selector) rather than the full candidates set, which requires a way to go from parameter-level operations to a candidates-level policy.
Limitations and Open Points
Refactoring needed before policies can be merged
TableCandidates: it should accept aLazyFrameso that the policy chain can stay directly in the lazy regime and never move to eager evaluation if not needed.validate_parameter_inputalso has to work on lazy frames, otherwise we cannot stay in the lazy data regime.comp_rep_columnsneeds refactoring to be dynamic -._make_input_scaler()makes use ofcomp_rep_columns, which we don't have at hand for the embeddings before runtime. Refactor to be dynamic.Thoughts on SequenceParameter
Notes on
SequenceParametergathered while testing it with policies:SequenceParameterprovides the definition and encoding; in theory,valueswill be infinite, but we can select suitable values by settingactive_valuesand use them as candidates. My suggestion:values- all potential candidates (can be infinite),active_values- recommendable values.CustomParameterwould require encoding the candidates immediately. But in a policy chain, we may want to do further operations on the sequence first, like random subsampling to heavily reduce the set, and only encode at the end. If the encoder for embeddings lived onCustomParameter, what wouldSequenceParametereven be for?valuesset is meant to stay infinite — we still need a way to express and store the possible values for that parameter.comp_rep_columns: Bigger refactoring is coming to make the column names dynamic.values; this is not feasible forSequenceParameter, asvaluesis always infinite in the current implementation --> Needs refactoring as well.To Be Discussed
is_finitebe treated forSequenceParameter? The underlyingParametercan be infinite, but by settingactive_valuesit can be made materializable, i.e. infinite != materializable. How should this be handled when working withSequenceParameterand Policies?