Skip to content

Add GP symmetries and Enable Auto-Symmetrized Kernels - #929

Open
Scienfitz wants to merge 11 commits into
feature/rgpe-with-kernel-overridefrom
feature/symmetric_kernels
Open

Scienfitz wants to merge 11 commits into
feature/rgpe-with-kernel-overridefrom
feature/symmetric_kernels

Conversation

@Scienfitz

@Scienfitz Scienfitz commented Oct 7, 2026 •

Copy link
Copy Markdown
Collaborator

Closes #622 and #620

This PR adds a symmetries attribute to GPs. In contrast to Recommender.symmetries (which triggers data augmentation) this triggers the construction of symmetric kernels in theGP if possible.

This is a complex topic and I don't consider this PR wisdoms last word, but tis a start and somewhat complete in terms of coverage.

TODO, Pending Until Architecture Agreement

  • finalize example and doc (will be done after principal architectural agreement)
  • update example pic
  • publish doc on fork

Proof of Concept

Here is the result from the expanded Symmetry example. It shows

  • Confirmed strongly positive impact
  • Edge over augmentation
  • Reduced impact in case of only adding data from the actual reduced searchspace (as expected)
image

-------------------- LLM based below --------------------

Mathematical Treatment of the Kernel Symmetries

Let $k$ denote the fully resolved kernel of the GP, i.e. the surrogate kernel including all parameter kernel overrides and the task kernel. All symmetries act on the normalized model inputs $x, x'$, i.e. after the GP's Normalize input transform. None of the constructions below add hyperparameters, all learnable parameters remain those of $k$.

A symmetrized kernel $\tilde k$ makes the model exactly invariant under a transformation $T$ if $\tilde k(Tx, x') = \tilde k(x, x')$ for all $x, x'$. The posterior mean $\mu(x) = m(x) + \tilde k(x, X) \alpha$ and the posterior variance only depend on the inputs via $\tilde k$ and the prior mean $m$. The posterior is therefore invariant for any hyperparameters, provided that $m$ is constant (enforced).

Permutation Symmetry

Equation (tied single sum). Let $\pi$ range over all $n!$ permutations of the $n$ positions of a PermutationSymmetry. $\pi x'$ permutes the positions of $x'$ in all groups in lockstep; parameters with several computational columns (e.g. one-hot encodings) are moved as blocks:

$$ k_{\text{perm}}(x, x') = \sum_{\pi} k(x, \pi x') $$

Validity. This is a valid kernel and invariant under all permutations if the base kernel is invariant under jointly permuting both inputs:

$$ k(\pi x, \pi x') = k(x, x') \quad \text{for all } \pi $$

In that case, $k_{\text{perm}}(x, x') = \tfrac{1}{n!} \sum_{\pi, \sigma} k(\pi x, \sigma x')$. BayBE guarantees the condition by sharing the hyperparameters of permuted positions:

  • Joint kernels spanning several positions (e.g. the default ARD Matérn): their per-dimension parameters (raw_lengthscale, raw_period_length) are tied entry-wise via a torch.nn.utils.parametrize parametrization (TiedEntries). The positions thus share a single learned lengthscale, e.g. $[\ell, \ell, \ell_z]$ for permuted $x_1, x_2$ and an unpermuted $z$.
  • Separate per-position kernels (e.g. equivalent override_kernels, or Matern(["x1"]) * Matern(["x2"])): the kernel of each position reuses the parameter objects of the corresponding first-position kernel.

Before fitting, a numerical check evaluates a copy of the tied kernel with random hyperparameters on random inputs and verifies $k(\pi x, \pi x') = k(x, x')$.

Limitations and whether they occur in BayBE.

Limitation Occurs in BayBE?
The base kernel must treat the permuted positions alike. Not for the default kernel or the presets (ARD lengthscales are tied). It occurs with user-specified kernels that model the positions differently (e.g. Matern(["x1"]) * RBF(["x2"])) or with a factory parameter_selector excluding only some permuted parameters. Both raise a ValueError before fitting. Different kernel overrides on permuted parameters are already rejected, since PermutationSymmetry requires equivalent parameters.
Per-dimension hyperparameters other than lengthscales and period lengths are not tied. Only for custom raw GPyTorch kernels with other per-dimension parameters. The numerical check catches them and raises an error before fitting.
Group size is limited to 5 positions ($5! = 120$ terms). Enforced at construction.
A tied lengthscale keeps one prior term per position, as untied lengthscales would. Always (by design). This slightly sharpens the prior relative to the data.
Re-attaching the lengthscale lower bound after tying relies on GPyTorch's private _constraints naming convention. Always. A test asserts that BoTorch still sees the bound.

Cost. $n!$ evaluations of the inner kernel per permutation symmetry with $n$
positions.

Mirror Symmetry

Equation (double sum). Let $m$ reflect the mirrored column at the normalized mirror point $c' = (c - \ell) / (u - \ell)$, i.e. replace its value $v$ by $2c' - v$, where $[\ell, u]$ are the scaling bounds of the parameter:

$$ k_{\text{mir}}(x, x') = k(x, x') + k(x, m x') + k(m x, x') + k(m x, m x') = \sum_{a \in {0, 1}} \sum_{ b \in {0, 1}} k(m^a x, m^b x') $$

Validity. This is a valid kernel and invariant under the reflection for any base kernel, since it equals the inner product of the feature maps $\phi(x) + \phi(m x)$. No hyperparameter tying is needed. For reflection-invariant base kernels (all kernels depending only on per-dimension distances, e.g. Matérn, RBF, RQ, periodic), the double sum equals twice the single sum $k(x, x') + k(x, m x')$. The learned output scale absorbs that factor.

Limitations and whether they occur in BayBE.

Limitation Occurs in BayBE?
The mirrored parameter must be numerical, i.e. a single computational column. Enforced by MirrorSymmetry itself.
Mirror images may lie outside the parameter range. Harmless: the kernel is defined on all real inputs.
None regarding the kernel. The double sum is also valid for linear, polynomial and raw GPyTorch kernels.

Cost. 4 evaluations of the inner kernel per mirror symmetry.

Dependency Symmetry

Equation (indicator gating). For $D$ dependency symmetries, let $a_d(x) \in {0, 1}$ indicate whether dependency $d$ is active at $x$, and let $P(x) = (a_1(x), \dots, a_D(x))$ be the activity pattern of $x$. For a pattern $P$, let $x_P$ denote $x$ with the affected columns of all dependencies inactive in $P$ set to a constant (the lower bound $0$ of the normalized range):

$$ k_{\text{dep}}(x, x') = \sum_{P \in {0, 1}^D} \mathbf{1}[P(x) = P] ; \mathbf{1}[P(x') = P] ; k(x_P, x'_P) $$

Validity. Each term is the product of the rank-one mask $\mathbf{1}[P(x) = P],\mathbf{1}[P(x') = P]$ and a kernel evaluated on transformed inputs, so the sum is positive semi-definite.

  • Inputs with an inactive dependency are compared as if their affected parameters were equal, so their values do not matter
  • Inputs with different patterns are uncorrelated, so the constant is never compared with an actual value
  • The causing parameter itself is never fixed, so several values that make a dependency inactive (e.g. off1, off2)
  • All patterns share one kernel and its hyperparameters, which works with any kernel.
  • The activity $a_d(x)$ is determined by matching the causing parameter's columns against the normalized computational representations of its values

Limitations and whether they occur in BayBE.

Limitation Occurs in BayBE?
Inputs with different activity patterns are uncorrelated, i.e. no information is shared between the regimes via the kernel. Always (deliberate modeling choice).
The causing parameter must be discrete. Enforced by DependencySymmetry itself.
The computational representation of the causing parameter must reveal the activity. Can occur for SubstanceParameter and CustomDiscreteParameter with decorrelation, which can map an active and an inactive value to identical encodings. Raises IncompatibleSearchSpaceError before fitting.
The condition must be evaluable on the causing parameter's values. Can occur, e.g. a ThresholdCondition on a categorical parameter. Raises an error before fitting via Condition.evaluate.
Degenerate conditions (all or no values active) are not rejected. Possible. The kernel then reduces to $k$ or always ignores the affected parameters. Validating this is deferred to the constraint refactoring.
The kernel matrix of each pattern is evaluated densely. Always. Irrelevant at BayBE's data sizes, since exact GP kernel matrices are dense anyway.

Cost. $2^D$ evaluations of the inner kernel for $D$ dependency symmetries (one per activity pattern, each masked to the matching pairs).

Overall Cost

With $G$ permutation symmetries with $n_1, \dots, n_G$ positions, $M$ mirror symmetries and $D$ dependency symmetries, each evaluation of the symmetric kernel requires

$$ E = 2^{D} \cdot \prod_{g=1}^{G} n_g! \cdot 4^{M} $$

evaluations of the resolved kernel $k$. This factor applies to building the kernel matrices, in every step of the hyperparameter optimization and at prediction time. The cost of the subsequent linear algebra is unchanged, since the number of training points stays the same. Data augmentation, in contrast, multiplies the number of training points.

Configuration $E$
One permutation with 2 / 3 / 5 positions 2 / 6 / 120
One / two mirror symmetries 4 / 16
One / three dependency symmetries 2 / 8
3-slot mixture (one permutation with 3 positions, 3 dependencies) 48
3-slot mixture plus one mirror symmetry 192

Combining Symmetries

Stack and Order

The symmetric kernel is built at the very end of _resolve_kernel, i.e. after all parameter kernel overrides and the task kernel have been applied. The resolved kernel $k$ is then wrapped, from the inside out:

$$ k_{\text{final}} = \underbrace{\text{Mirror}_M \circ \dots \circ \text{Mirror}_1}_{\text{reflection sums}} \circ \underbrace{\text{Perm}_G \circ \dots \circ \text{Perm}_1}_{\text{permutation sums}} \circ \underbrace{\text{Gate}}_{\text{dependencies}} (k) $$

Before wrapping, the hyperparameters of all permuted positions in $k$ are tied, and the joint permutation invariance of $k$ is verified numerically.

  1. Dependency gate (innermost). It needs the raw values of the affected and causing
    columns to determine the activity patterns and to fix inactive affected columns.
  2. Permutation sums. They require the kernel they wrap to be invariant under jointly
    permuting both inputs. The gated kernel satisfies this if $k$ is tied and the
    dependencies are closed under the permutation (see below). The permutation then only
    renumbers the dependencies, i.e.
    $k_{\text{dep}}(\pi x, \pi x') = k_{\text{dep}}(x, x')$. In the reverse order, the gate
    would determine the activity patterns on unpermuted inputs and break the permutation
    invariance.
  3. Mirror sums (outermost). They are valid for any inner kernel. Mirrored parameters
    are neither permuted by another symmetry nor causing parameters. The reflections
    therefore commute with the permutations and never change an activity pattern, so all
    inner invariances are preserved.

Because the posterior also depends on the prior mean, only constant means (or task-wise constant means in transfer learning) are accepted. In RGPE mode, each inner GP of the ensemble builds the same stack on its own search space.

Allowed Combinations (LLM based)

Each wrapper always yields a valid kernel, but the invariances only combine if the roles of the involved parameters are compatible. A parameter can be permuted, mirrored, causing (decides whether a dependency is active) or affected (only matters if a dependency is active). validate_symmetries enforces the following rules at construction time:

Combination Allowed Reason
Parameter permuted or mirrored by several symmetries No The transformations interfere, so one invariance is lost under the other.
Mirrored causing parameter No The reflection can change whether the dependency is active, so an inactive affected value suddenly matters.
Causing parameter affected by another dependency (chain) No An inactive parameter would still decide whether another one matters.
Permuted or mirrored affected parameter Yes When the dependency is inactive, the affected value is fixed anyway.
Parameter in several dependencies Yes It is fixed whenever any of its dependencies is inactive.
Dependencies not closed under a permutation No The permutation would change the gated kernel, so the permutation sum is not invariant.

Closure means that renaming the causing and affected parameters of any dependency according to any permutation yields another existing dependency with an identical condition (compared via ==). Dependencies that touch no permuted parameter satisfy this automatically. Permutation groups are additionally limited to 5 positions.

Slot-based mixtures. The closure rule enables the main use case of combining symmetries. A mixture modeled with slots permutes the substance labels and amounts of all slots in lockstep, and each label only matters if the amount of its slot is positive:

  • PermutationSymmetry([["Solvent_1", "Solvent_2", "Solvent_3"], ["Fraction_1", "Fraction_2", "Fraction_3"]])
  • DependencySymmetry("Fraction_i", ThresholdCondition(0, ">"), ["Solvent_i"]) for
    $i = 1, 2, 3$

Permuting the slots maps the dependency of slot $i$ onto the dependency of slot $\pi(i)$, so the set of dependencies is unchanged and all invariances combine. This is exactly what DiscretePermutationInvarianceConstraint.to_symmetry() and DiscreteDependenciesConstraint.to_symmetries() produce for such a setup, provided all slots use identical conditions. Mixtures whose slots use equivalent but differently specified conditions are rejected. A 3-slot mixture costs $2^3 \cdot 3! = 48$ kernel evaluations.

@Scienfitz Scienfitz added this to the 0.16.0 milestone Oct 7, 2026
@Scienfitz Scienfitz self-assigned this Oct 7, 2026
Copilot AI balanced review requested due to automatic review settings October 7, 2026 18:01
@Scienfitz
Scienfitz requested a review from AdrianSosic as a code owner October 7, 2026 18:01
@Scienfitz
Scienfitz added this pull request to stack #905 October 7, 2026 18:01
@Scienfitz
Scienfitz requested a review from AVHopp as a code owner October 7, 2026 18:01
@Scienfitz Scienfitz added the new feature New functionality label Oct 7, 2026

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Changes recommended

Unsupported custom symmetry subclasses are silently ignored instead of being rejected or enforced.

1 open finding
What changed in this PR

Adds exact symmetry-aware Gaussian process kernels while preserving recommender-level data augmentation.

Changes:

  • Implements permutation-, mirror-, and dependency-invariant kernel wrappers.
  • Adds GP symmetry validation, parameter tying, RGPE propagation, and tests.
  • Updates documentation, examples, constraints, and search-space lookup behavior.
File Description
baybe/​surrogates/​gaussian_process/​_symmetry.py Constructs and validates invariant kernels.
baybe/​surrogates/​gaussian_process/​components/​_gpytorch.py Adds GPyTorch symmetry wrappers and parameter tying.
baybe/​surrogates/​gaussian_process/​core.py Exposes GP symmetries and integrates kernel construction.
baybe/​surrogates/​gaussian_process/​_override/​kernel.py Adds kernel-tree traversal.
baybe/​surrogates/​gaussian_process/​_override/​__init__.py Exports kernel-tree traversal.
baybe/​symmetries/​base.py Documents invariant-kernel usage.
baybe/​symmetries/​dependency.py Validates dependency roles and clarifies discretization.
baybe/​constraints/​discrete.py Rejects self-dependent parameters.
baybe/​searchspace/​core.py Generalizes name collections.
baybe/​searchspace/​discrete.py Rejects bare-string name queries.
baybe/​searchspace/​continuous.py Rejects bare-string name queries.
baybe/​recommenders/​pure/​bayesian/​base.py Distinguishes augmentation from kernel symmetry.
tests/​test_invariant_kernels.py Tests invariance, validity, and posterior behavior.
tests/​validation/​test_symmetry_validation.py Tests invalid GP symmetry configurations.
tests/​validation/​test_constraint_validation.py Tests self-dependency rejection.
tests/​test_searchspace.py Tests exact parameter-name lookup.
tests/​test_rgpe_surrogate.py Tests symmetry propagation into RGPE models.
tests/​test_parameter_kernel_overrides.py Tests symmetry composition with overrides.
tests/​test_iterations.py Adds end-to-end symmetric GP coverage.
tests/​test_gp.py Tests preset symmetry configuration.
tests/​hypothesis_strategies/​surrogates.py Generates symmetry-enabled GP configurations.
examples/​Symmetries/​permutation.py Compares constraints, augmentation, and invariant kernels.
docs/​components/​symmetries.md Documents invariant-kernel mathematics and limitations.
docs/​components/​surrogates.md Links surrogate users to symmetry documentation.
docs/​components/​recommenders.md Documents the kernel-based alternative.
docs/​components/​parameters.md Documents override compatibility.
docs/​components/​constraints.md Documents self-dependency restrictions.
CHANGELOG.md Records API and feature changes.

🧠 Review effort: Balanced


💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment on lines +68 to +70
permutations = [s for s in symmetries if isinstance(s, PermutationSymmetry)]
mirrors = [s for s in symmetries if isinstance(s, MirrorSymmetry)]
dependencies = [s for s in symmetries if isinstance(s, DependencySymmetry)]
A single string passed as names was matched via substring membership, so looking
up "n1_not_discrete" also returned a parameter named "n1". DependencySymmetry
passed its causing parameter name this way and could thus validate the wrong
parameter. Strings are now rejected with a ValueError and the affected call site
passes a tuple.
Adds iter_gpytorch_kernel_tree, which yields every kernel of a GPyTorch kernel tree
with the input columns it acts on, and rebuilds get_active_dimensions on top of it.
The results are unchanged for all kernels built by BayBE, but nested active_dims and
arbitrary composite kernels are now handled correctly.
Replaces the blanket ban on symmetries sharing parameters with rules based on the
roles of the parameters. In particular, dependencies are now allowed on permuted
parameters as long as they are closed under the permutations, which enables
slot-based mixtures.
@Scienfitz
Scienfitz force-pushed the feature/symmetric_kernels branch from a117378 to 545f3e7 Compare October 7, 2026 18:37

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

new feature New functionality

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Add Invariant GP Kernels

2 participants