24August2026
14:00 Doctoral defense room 85 of IC2
Topic on
Towards Robust Visual Odometry under Camera Failures and Diverse Motion Profiles: From Learned Frontends to Geometry-Informed Dense–Sparse Fusion
Student
Hudson Martins Silva Bruno
Advisor / Teacher
Esther Luna Colombini
Brief summary
Visual Odometry (VO) is fundamental for autonomous systems; however, classical geometry-based systems remain highly fragile when deployed in unstructured environments or subject to camera failures. This thesis investigates the integration of deep learning into VO pipelines to overcome these vulnerabilities. The research initially evaluates hybrid frameworks that incorporate learning-based feature extractors and attention-based matching methods in traditional backends. These fundamental experiments reveal a critical trade-off: while deep neural networks significantly increase the robustness of local tracking against camera failures, they often struggle to maintain the long-term global consistency inherent in classical methods. To quantify these vulnerabilities and assess the robustness of the systems, this work introduces the QueensCAMP dataset. This new benchmark systematically injects standardized camera failures (such as lens condensation, dirt, and breakage) into challenging real-world trajectories. Evaluations using this dataset confirm that, although end-to-end deep learning models exhibit superior continuous tracking capabilities compared to classical geometry-based SLAM, they remain susceptible to uncorrected errors when generalizing across distinct domains. Based on these insights, the research explores the capabilities of Vision Transformers (ViTs) for end-to-end VO. To align semantic attention mechanisms with precise spatial tracking, a novel geometric curriculum learning strategy is proposed, using self-supervised homography estimation as a pre-training proxy task. Finally, these combined investigations culminate in the development of Deep Direct-Indirect Visual Odometry (DDI-VO), a novel monocular architecture. DDI-VO fundamentally unifies the dense direct and sparse indirect VO paradigms, merging the global contextual perception of a homography-calibrated ViT with the structural rigidity of sparse geometric anchors through a cross-attention mechanism. Evaluations were performed on heterogeneous benchmarks (autonomous driving, high-speed flight, and irregular indoor scenarios with a handheld camera) to demonstrate the architecture's robustness. While DDI-VO did not achieve the lowest error in all benchmarks, it consistently performed among the best in the automotive, aerial, and indoor domains, decisively exceeding both baselines under camera failure conditions in QueensCAMP, establishing itself as a robust, generalist framework for accurate camera motion estimation in diverse environments, distinct motion profiles, and under camera failure conditions.
Examination Board
Headlines:
Esther Luna Colombini IC / UNICAMP
Gustavo Teodoro Laureano INF/UFG
Valdir Grassi Junior EESC/USP
André Oliveira Françani EESC/USP
Alisson Vasconcelos de Brito CI/UFPB
Substitutes:
Gabriel Capiteli Bertocco IC / UNICAMP
Fernando dos Santos Osório ICMC / USP
Flavio Tonidandel Socialdroids Robotics
Institute of Computing
Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognizing you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.