Skip to content

v3.6.6: estimator, numerical and CLI fixes - #749

Merged
yzhao062 merged 17 commits into
masterfrom
development
Sep 17, 2026
Merged

yzhao062 merged 17 commits into
masterfrom
development

Conversation

@yzhao062

Copy link
Copy Markdown
Owner

PyOD 3.6.6

A maintenance release focused on estimator compatibility, numerical correctness, and documentation. No new detector or dependency is introduced.

Bug fixes

  • LOCI: preserve the constructor parameter k for cloning and use contamination to derive the fitted threshold and binary labels. k still controls the radius-scan early stop. #707, fixes #194.
  • Pearson correlation: compute all row pairs in the unweighted pearsonr_mat utility when there are more rows than features; previously some correlations incorrectly remained 1. #731.
  • Deep-learning base class: return the fitted estimator from BaseDeepLearningDetector.fit. #740.
  • VAE: preserve logvar_clip as supplied to the constructor so sklearn.clone works, while validating the training value separately. #738.
  • Sampling: retain a fractional subset_size across fits instead of replacing the constructor argument with the first computed sample count. #737.
  • Kernel PCA: forward the requested n_components to the underlying estimator. #730.
  • KNN: allow supported brute-force metrics such as cosine without constructing an incompatible fallback BallTree. #729.
  • DeepSVDD: suppress epoch-loss logging when verbose=0. #735.
  • CLI: avoid falsely reporting Claude Code as installed merely because pyod install skill created its skill directory. Detection checks the executable or user configuration instead. #723, fixes #715.

Documentation and tests

  • Remove the unsupported sample_weight entry from GMM.fit documentation. #741.
  • Describe score_to_label as returning binary labels, not probabilities. #733.
  • Repair KNN, LOF, and LUNAR reference links. #732.
  • Clarify the ECOD/COPOD relationship. #725, addresses #655.
  • Clarify LOCI's early-stop parameter, restore weighted Pearson range assertions, and seed the new rectangular-input regression fixture.
  • Add regression coverage for estimator contracts, repeated fits, component forwarding, cosine distance, logging, contamination and rectangular correlations; make CLI detection tests independent of the host's installed executables.
  • Retain the repository's existing CITATION.cff on master, alongside contributor-workflow ignore rules and the todo/ drop box.

Behavior notes

LOCI's binary labels now follow the documented contamination-based threshold rather than a fixed k cutoff. Raw LOCI scores for a fixed k are unchanged by this fix; tied scores can still make the labeled fraction differ from the requested contamination.

Kernel PCA scores can change when n_components is specified because the limit is now honored. Sampling refits with a fractional subset_size now use the current training-set size. Previously uncomputed Pearson row pairs now receive their actual correlation values.

The ROD duplicate-weighting proposal #746 is not included.

Contributors

Thanks to @Mohit-Ak, @VenishPaneliya, @Iams4kura, @busysofa15, @godarrenw, and @dishasharma23-prog.

Full changelog: v3.6.5...v3.6.6

busysofa15 and others added 17 commits August 6, 2026 19:53
FIX: Prevent false-positive Claude Code detection in pyod info (#715)
* Respect verbose in DeepSVDD training output

DeepSVDD documents `verbose` ("Verbosity mode", default 1) and stores it
on the estimator, but the training loop prints the per-epoch loss
unconditionally. `self.verbose` is never read, so `DeepSVDD(verbose=0)`
still writes one line per epoch - 100 lines at the default `epochs`, for
every fit in a benchmark loop.

With epochs=3:

    verbose=0  ->  3 epoch lines printed
    verbose=1  ->  3 epoch lines printed

The print is now gated on `self.verbose == 1`, which is how AnoGAN and
LUNAR already gate their per-epoch output. The default is 1, so nothing
changes unless the caller asked for silence.

* Print epoch logs for any non-zero verbose

Gating on `verbose == 1` made `DeepSVDD(verbose=2)` quieter than
`verbose=1`, which inverts the usual meaning of the setting. Suppress the
per-epoch line only when verbosity is 0, and cover mode 2 in the test.
`GMM.fit` is `fit(self, X, y=None)`, but its docstring documents a third
argument:

    sample_weight : array-like, shape (n_samples,)
        Per-sample weights. Rescale C per sample. Higher weights
        force the classifier to put more emphasis on these points.

Passing it raises `TypeError: GMM.fit() got an unexpected keyword
argument 'sample_weight'`.

The text is copied from `OCSVM.fit`, which does take the argument - "C"
is the SVM penalty parameter and has no counterpart in a Gaussian
mixture. It also cannot simply be forwarded: `sklearn.mixture
.GaussianMixture.fit` is `fit(self, X, y)` and has no `sample_weight`.

Only the stale documentation is removed; OCSVM is untouched.
Signed-off-by: Zhu yizhang <95731595+godarrenw@users.noreply.github.com>
Signed-off-by: Zhu yizhang <95731595+godarrenw@users.noreply.github.com>
`BaseDetector.fit` documents its contract as

    Returns
    -------
    self : object

but the deep learning base ends `fit` on `_process_decision_scores()`
without returning, so every detector built on it returns None:

    VAE().fit(X)          -> None
    AutoEncoder().fit(X)  -> None
    DeepSVDD().fit(X)     -> DeepSVDD      (defines its own fit)
    IForest().fit(X)      -> IForest

That breaks the chained form these estimators otherwise support -
`clf.fit(X).predict(X)` raises AttributeError on None - and any sklearn
utility that relies on `fit` returning the estimator.

Adds the missing return and the Returns section the base class already
documents.
`VAE.__init__` stored the validator's return value rather than the
argument:

    self.logvar_clip = self._validate_logvar_clip(logvar_clip)

and `_validate_logvar_clip` ends with `return float(lower), float(upper)`,
so the attribute is a freshly built tuple. sklearn's `clone()` checks
that the constructor left its parameters untouched using an identity
comparison, so this fails outright:

    RuntimeError: Cannot clone object VAE(...), as the constructor either
    does not set or modifies parameter logvar_clip

which means `VAE` cannot be used with anything that clones - GridSearchCV,
cross_val_score, or a Pipeline being refit.

The argument is now stored as given. Validation still happens in
`__init__`, so an invalid `logvar_clip` is still rejected at construction
time and the existing config tests are unaffected. The normalised floats
move to `self.logvar_clip_`, derived in `build_model`, which is what the
model and the training loss now read.
`Sampling.fit` resolves a fractional `subset_size` by writing the
absolute count back onto the estimator:

    self.subset_size = int(self.subset_size * n_samples)

so the constructor argument is destroyed on the first fit. Two
consequences:

* refitting on a differently sized dataset silently reuses the stale
  count. With `subset_size=0.2`, fitting on 100 rows samples 20, and a
  refit on 400 rows samples 20 again instead of 80;
* `get_params()["subset_size"]` returns 20 rather than the 0.2 that was
  passed, so a `clone()` taken after a fit is not the estimator the user
  configured. sklearn requires `__init__` parameters to be left as they
  were given.

The resolved value is now a local, and `self.subset_size` is left alone.

While rewriting the surrounding lines, the `ValueError` for an
out-of-range fraction was missing its `%` operand, so it reported the
literal "subset_size=%r must be between 0.0 and 1.0". It now interpolates
the value, like the integer branch just above it already does.
`PyODKernelPCA.__init__` takes `n_components` but does not pass it to
`super().__init__()`, so the inner `KernelPCA` keeps its default of
`None`. It is the only one of the sixteen parameters that is not
forwarded.

`KPCA` exposes the argument publicly and documents it, so
`KPCA(n_components=3)` silently computes every component instead of 3:

    KPCA(n_components=3).fit(X)   # X is (420, 8)
    inner n_components        -> None
    components computed       -> 419
    transform(X).shape        -> (420, 419)

The scores are unaffected, because `decision_function` slices
`x_transformed[:, :n_selected_components_]` afterwards and the surplus
components contribute nothing. What it costs is time and memory: at
n=2000 with `n_components=3` the fit goes from 8.60s to 0.50s and the
transform matrix from 2000x1999 (32.0 MB) to 2000x3, because with the
inner value left at `None` sklearn always takes the dense `eigh` path.

`remove_zero_eig=False` is also ineffective today, since sklearn forces
zero-eigenvalue removal when `n_components is None`.

Checked `decision_scores_` before and after across seven configurations
(`n_components` 1/3/100, `n_selected_components`, `eigen_solver="arpack"`,
`remove_zero_eig=True`, `kernel="linear"`, and defaults): the largest
difference is 6.9e-14.
LOCI stored the constructor argument k directly into self.threshold_ and
never called _process_decision_scores(), so contamination was accepted but
completely ignored when deriving labels_. Store k on the instance, use it
in the score loop where it belongs, and let the base class compute
threshold_ from contamination like every other detector.

Fixes #194
`pearsonr_mat` builds an `(n_row, n_row)` matrix and fills it pairwise,
but the unweighted branch runs its outer loop over `n_col`:

    if w is not None:
        for cx in range(n_row):
            ...
    else:
        for cx in range(n_col):

The weighted branch directly above uses `n_row`, and the docstring
describes the result as "Row-wise pearson score matrix" of shape
(n_samples, n_samples).

When there are fewer columns than rows, every pair with `cx >= n_col` is
skipped and those cells keep the value the matrix was initialised with,
so the function reports a correlation of exactly 1.0 for pairs it never
looked at. On a (50, 4) matrix, 2070 of 2500 entries are wrong, with
errors up to 2.0.

Output is unchanged whenever `n_col >= n_row - 1`, and the weighted
branch is untouched. `n_col` has no remaining reader, so it goes too.

The existing test uses a (10, 20) matrix, which is in the unaffected
range, and only checks the shape; the added test compares every entry
against `scipy.stats.pearsonr` on an (8, 3) matrix.
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 17, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-17T01:34:09.917448Z 97fa2b4 PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@coveralls

coveralls commented Sep 17, 2026 •

Copy link
Copy Markdown

Coverage Report for CI Build 35170878889

Coverage increased (+0.05%) to 92.625%

Details

  • Coverage increased (+0.05%) from the base build.
  • Patch coverage: 1 uncovered change across 1 file (131 of 132 lines covered, 99.24%).
  • No coverage regressions found.

Uncovered Changes

File Changed Covered %
pyod/cli.py 2 1 50.0%
Total (17 files) 132 131 99.24%

Coverage Regressions

No coverage regressions found.


Coverage Stats

Coverage Status
Relevant Lines: 20813
Covered Lines: 19278
Line Coverage: 92.62%
Coverage Strength: 10.19 hits per line

💛 - Coveralls

@yzhao062
yzhao062 merged commit 3a71f47 into master Sep 17, 2026
26 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

pyod info detects Claude Code from a directory that pyod install skill creates Contamination parameter in LOCI

8 participants