Image and Video Coding for Humans and Machines
Images and videos are increasingly consumed not only by humans but also by recognition models. In collaboration with Yui Tatsumi (Watanabe Lab), this project studies coding frameworks that serve both machine vision and human viewing.
- Explicit residual-based scalable coding (FR-ICMH / PR-ICMH) improves the coding efficiency and interpretability of scalable coding for humans and machines, with up to 29.57% BD-rate savings. (IEEE MMSP 2025)
- Seed selection picks the best random seed from early reverse-diffusion outputs, improving human-oriented reconstruction without extra bitrate. (IEEE GCCE 2025, Oral Presentation Award)
- Training-free adaptive quantization enables continuous variable-rate control for image coding for machines, with up to 11.07% BD-rate savings. (IEEE ICCE 2026)
- Scalable coding for humans and machines extended to video via keyframe interpolation. (MIRU 2025; MIRU 2026, Interactive Presentation Award)