Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

NYBG Herbarium Image Classification

I first worked on this problem in 2024 through Break Through Tech AI and a New York Botanical Garden Kaggle challenge. The task was to classify collection images into ten categories, including pressed specimens, live plants, microscope slides, and illustrations.

The original project lived mostly in a notebook. I returned to the problem in 2026 with fresh eyes and rebuilt it as a Python package so I could rerun the model, question the original choices, and take a closer look at its mistakes. This repository is that rebuild, not the original competition submission.

Rosa palustris herbarium specimen from The New York Botanical Garden

Rosa palustris specimen from The New York Botanical Garden's William and Lynda Steere Herbarium. Image courtesy of The New York Botanical Garden.

Results

I trained a ResNet50 model with ImageNet weights on 81,946 images and evaluated it on a separate set of 10,244 images. For this run, I kept the ResNet50 backbone frozen and trained the classification head without augmentation.

Metric Result
Validation accuracy 95.69%
Macro F1 95.71%
Incorrect predictions 442 / 10,244

Normalized validation confusion matrix

The model had the most trouble separating mixed pressed specimens from ordinary pressed specimens. It also performed differently across image sources. For example, accuracy for source MPU was 81.6% on 76 images, well below the overall accuracy. That gap is one reason I would not treat 95.69% as a complete measure of how the model would perform on a new museum collection.

Training and validation accuracy across eight epochs

Ten lowest source-level validation accuracy values

Source groups have different sample sizes, so I use this breakdown to find places to investigate rather than to rank the contributing collections.

These are validation results from this rebuild. I am not claiming a historical Kaggle score.

Follow-up experiments

I ran more tests after rebuilding the baseline. The biggest difference showed up when I removed one image source from training at a time. Accuracy was 96.56% on the regular split and 81.86% with each source held out.

Experiment Result
ResNet50 feature probe, ordinary split 96.56% accuracy
Same probe protocol, each source held out 81.86% weighted accuracy
Source metadata only, without image pixels 64.45% accuracy
Likely near-identical cross-split pairs 7
Entropy-based OOD detection on MNIST 0.516 AUROC

Random-split, source-held-out, and source-only accuracy

Frozen-feature learning curves

ResNet50 and EfficientNetV2-B0 were nearly tied across three runs: 96.40% and 96.33%. I put the rest of the results in the experiment notes.

Returning to the problem

Why did I freeze ResNet50?

I wanted a clear transfer-learning baseline before changing the entire model. The frozen ImageNet backbone supplied general visual features while keeping the training cost manageable and limiting the number of parameters that could overfit. It also made the experiment easier to interpret: this result measures what a new classification head can learn from fixed ResNet50 features.

Why did validation accuracy peak at epoch four?

Training accuracy continued to rise, but validation accuracy moved up and down after the fourth epoch. That is a sign that the head was starting to fit details of the training split that did not consistently help on validation data. Epoch six was nearly tied, so I do not treat epoch four as a magic number; I kept the best checkpoint and used early stopping rather than the final epoch.

Why were the pressed-specimen classes confused?

Mixed and ordinary pressed specimens are visually close, and the distinction can depend on composition and context rather than one obvious object. Resizing every image to a square may also reduce useful layout information. Of the 442 errors, 103 mixed specimens were predicted as ordinary and 120 ordinary specimens were predicted as mixed.

What might explain the differences between image sources?

The result shows an association, not a cause. Different collections may use different backgrounds, mounts, labels, cameras, lighting, or digitization workflows. The source groups also have different sample sizes. A model can learn those visual signatures along with the intended categories, so a random split may look better than performance on a genuinely unfamiliar collection.

What experiment did I run next?

I held out each image source in turn and tried two other resize methods. Source-held-out accuracy fell to 81.86%, with especially large drops for sources L and K. Center cropping did not help. I would test source-aware training next.

What did I rebuild?

I rebuilt nearly every working part represented in this repository: metadata validation, image loading and preprocessing, the model and training loop, evaluation and prediction commands, uncertainty and provenance output, tests, CI, and the result figures. The original competition experience was collaborative; this repository documents my later end-to-end return to the problem.

What is in the repository

  • reusable commands for training, evaluation, and prediction;
  • checks for metadata, labels, image paths, and submission shape;
  • per-class and per-source evaluation;
  • confidence and entropy output for reviewing uncertain predictions;
  • source-held-out, learning-curve, resize, architecture, and OOD probes;
  • exact and perceptual duplicate review without automatic deletion;
  • hashes and runtime information for tracking each experiment; and
  • tests and a GitHub Actions workflow.

Running it

The competition images and metadata are not included in this repository. They must be downloaded through an authorized Kaggle account.

python -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e '.[ml,data,dev]'

python -c 'import kagglehub; kagglehub.login()'
nybg-download --output-dir data/raw

Once the files are extracted, the training command takes explicit paths to the CSV files and image directories:

nybg-train \
  --train-csv path/to/train.csv \
  --validation-csv path/to/validation.csv \
  --train-images path/to/train-images \
  --validation-images path/to/validation-images \
  --output-dir artifacts/resnet50

Credit

The 2024 competition project was collaborative. This later rebuild is my own return to the problem and should not be read as a claim that the original project was individual work.

The figures in this README can be regenerated with:

python scripts/generate_readme_figures.py

About

Image classification experiments using the NYBG herbarium dataset

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages