Fast and Easy Infinite Neural Networks in Python
-
Updated
Mar 1, 2024 - Jupyter Notebook
Fast and Easy Infinite Neural Networks in Python
CVPR 2024-Improved Implicit Neural Representation with Fourier Reparameterized Training
ICML2025-Inductive Gradient Adjustment for Spectral Bias in Implicit Neural Representations
Existing literature about training-data analysis.
A unified framework for attributing model behavior to model components, training data, and training dynamics.
Official repository for "FOCUS: First Order Concentrated Updating Scheme"
Code for "What Happens During the Loss Plateau? Understanding Abrupt Learning in Transformers" (NeurIPS 2025)
Code for 'Towards a Theoretical Understanding of the 'Reversal Curse' via Training Dynamics'
Source code for <Probability Consistency in Large Language Models: Theoretical Foundations Meet Empirical Discrepancies>
Code for "Effect of equivariance on training dynamics"
Official repository for the EMNLP 2024 paper "How Hard is this Test Set? NLI Characterization by Exploiting Training Dynamics"
A two-parameter Weibull lens on transformer weights — diagnose weight-magnitude distributions (shape k, scale λ) across 7 model families, and explain how λ evolves under AdamW training. Library (npm-weibull-py) + benchmark database + companion code for arXiv:2605.18898 and 2606.19367.
TMLR 2026 | Mechanistic interpretability: attention-head binding (EB*) as a marker of concept emergence. 7 models, 5 architectures (Pythia 160M–2.8B, OLMo-1B, CRFM GPT-2, SmolLM3-3B, Qwen2.5-1.5B), 41 terms.
Cross-Family Convergence of Neural Network Weight Skeletons. Companion to Zenodo paper (10.5281/zenodo.19652706).
Code and data for: Three Phases of Expert Routing — How Load Balance Evolves During MoE Training
Effective rank, RankMe, E1, CKA and anisotropy on transformer hidden states are determined by one direction. The exact identity, and the attention sink behind it.
A plug-in debugger and visualizer for RL reward functions. Detects reward hacking, tracks training health, and renders a live terminal dashboard.
[Ongoing] Post-training dynamics for SFT/RL checkpoint trajectories, concept directions, and structural analysis.
Code for "Abrupt Learning in Transformers: A Case Study on Matrix Completion" (NeurIPS 2024)
Atomic benchmark suite showing drift can act as an early warning before direct symmetry detection in gradual-breaking regimes, with reversal controls, finite-budget sensitivity tests, and exact alarm-time validation.
Add a description, image, and links to the training-dynamics topic page so that developers can more easily learn about it.
To associate your repository with the training-dynamics topic, visit your repo's landing page and select "manage topics."