I first worked on this problem in 2024 through Break Through Tech AI and a New York Botanical Garden Kaggle challenge. The task was to classify collection images into ten categories, including pressed specimens, live plants, microscope slides, and illustrations.
The original project lived mostly in a notebook. I returned to the problem in 2026 with fresh eyes and rebuilt it as a Python package so I could rerun the model, question the original choices, and take a closer look at its mistakes. This repository is that rebuild, not the original competition submission.
Rosa palustris specimen from The New York Botanical Garden's William and Lynda Steere Herbarium. Image courtesy of The New York Botanical Garden.
I trained a ResNet50 model with ImageNet weights on 81,946 images and evaluated it on a separate set of 10,244 images. For this run, I kept the ResNet50 backbone frozen and trained the classification head without augmentation.
| Metric | Result |
|---|---|
| Validation accuracy | 95.69% |
| Macro F1 | 95.71% |
| Incorrect predictions | 442 / 10,244 |
The model had the most trouble separating mixed pressed specimens from ordinary
pressed specimens. It also performed differently across image sources. For
example, accuracy for source MPU was 81.6% on 76 images, well below the overall
accuracy. That gap is one reason I would not treat 95.69% as a complete measure
of how the model would perform on a new museum collection.
Source groups have different sample sizes, so I use this breakdown to find places to investigate rather than to rank the contributing collections.
These are validation results from this rebuild. I am not claiming a historical Kaggle score.
I ran more tests after rebuilding the baseline. The biggest difference showed up when I removed one image source from training at a time. Accuracy was 96.56% on the regular split and 81.86% with each source held out.
| Experiment | Result |
|---|---|
| ResNet50 feature probe, ordinary split | 96.56% accuracy |
| Same probe protocol, each source held out | 81.86% weighted accuracy |
| Source metadata only, without image pixels | 64.45% accuracy |
| Likely near-identical cross-split pairs | 7 |
| Entropy-based OOD detection on MNIST | 0.516 AUROC |
ResNet50 and EfficientNetV2-B0 were nearly tied across three runs: 96.40% and 96.33%. I put the rest of the results in the experiment notes.
Why did I freeze ResNet50?
I wanted a clear transfer-learning baseline before changing the entire model. The frozen ImageNet backbone supplied general visual features while keeping the training cost manageable and limiting the number of parameters that could overfit. It also made the experiment easier to interpret: this result measures what a new classification head can learn from fixed ResNet50 features.
Why did validation accuracy peak at epoch four?
Training accuracy continued to rise, but validation accuracy moved up and down after the fourth epoch. That is a sign that the head was starting to fit details of the training split that did not consistently help on validation data. Epoch six was nearly tied, so I do not treat epoch four as a magic number; I kept the best checkpoint and used early stopping rather than the final epoch.
Why were the pressed-specimen classes confused?
Mixed and ordinary pressed specimens are visually close, and the distinction can depend on composition and context rather than one obvious object. Resizing every image to a square may also reduce useful layout information. Of the 442 errors, 103 mixed specimens were predicted as ordinary and 120 ordinary specimens were predicted as mixed.
What might explain the differences between image sources?
The result shows an association, not a cause. Different collections may use different backgrounds, mounts, labels, cameras, lighting, or digitization workflows. The source groups also have different sample sizes. A model can learn those visual signatures along with the intended categories, so a random split may look better than performance on a genuinely unfamiliar collection.
What experiment did I run next?
I held out each image source in turn and tried two other resize methods.
Source-held-out accuracy fell to 81.86%, with especially large drops for
sources L and K. Center cropping did not help. I would test source-aware
training next.
What did I rebuild?
I rebuilt nearly every working part represented in this repository: metadata validation, image loading and preprocessing, the model and training loop, evaluation and prediction commands, uncertainty and provenance output, tests, CI, and the result figures. The original competition experience was collaborative; this repository documents my later end-to-end return to the problem.
- reusable commands for training, evaluation, and prediction;
- checks for metadata, labels, image paths, and submission shape;
- per-class and per-source evaluation;
- confidence and entropy output for reviewing uncertain predictions;
- source-held-out, learning-curve, resize, architecture, and OOD probes;
- exact and perceptual duplicate review without automatic deletion;
- hashes and runtime information for tracking each experiment; and
- tests and a GitHub Actions workflow.
The competition images and metadata are not included in this repository. They must be downloaded through an authorized Kaggle account.
python -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -e '.[ml,data,dev]'
python -c 'import kagglehub; kagglehub.login()'
nybg-download --output-dir data/rawOnce the files are extracted, the training command takes explicit paths to the CSV files and image directories:
nybg-train \
--train-csv path/to/train.csv \
--validation-csv path/to/validation.csv \
--train-images path/to/train-images \
--validation-images path/to/validation-images \
--output-dir artifacts/resnet50The 2024 competition project was collaborative. This later rebuild is my own return to the problem and should not be read as a claim that the original project was individual work.
The figures in this README can be regenerated with:
python scripts/generate_readme_figures.py