I work on computer vision and its path into the physical world: helping machines see, understand, and eventually act in everyday environments.
- 👁️ Visual perception: segmentation, detection, and open-vocabulary recognition
- 🧠 Vision-language models: multimodal understanding, grounding, and efficient deployment
- 🧪 Data-centric & generative AI: synthetic data, automatic annotation, and data selection
- 🤖 Embodied AI: 3D scenes, simulation, and robot navigation
My earlier work was in biomedical image segmentation (semi-supervised and continual learning). Since then I've gradually moved toward systems that perceive, reason, and act.
I also like making small tools for problems I run into along the way.
| Project | Description |
|---|---|
| LabelEditor for VS Code | Image annotation in VS Code, with SAM-assisted masks |
| MiniMax H3 Bot | Multi-GPU image-to-video service for ComfyUI and Telegram |
| Sidebrowser | A compact side-panel browser for Windows |
| Neon Postgres Sync | Two-way sync between local files and Postgres in VS Code |
| Claude Telegram Bot | A Claude chatbot for Telegram |
Website · Email · Instagram · YouTube
✨ AI-generated profile summary

