My research asks what it would take to build intelligence trustworthy enough to act in high stakes visual domains, the kind, like medical imaging, where a confident mistake carries real consequences. I work across the layers this demands, perception, grounding, reasoning, and self improvement, and I am actively advancing each. Below is how they connect, and the direction I am reaching toward next.
Teaching models to perceive complex visual domains at scale, and measuring honestly what they still cannot see. Quilt-1M, MedicalNarratives, MedBLINK, MsCAMIL.
Binding language to precise regions in space and time, the bridge from seeing to locating. Quilt-LLaVA, SVG2.
Composing perception and grounding into explainable, iterative decision policies that can rival experts. PathFinder.
Learning rewards where no ground truth reference exists, so systems can improve on open ended tasks. When Rubrics Fail.
The frontier I am most drawn to: estimating the physical properties of soft, deformable tissue from observation, so a system can predict how it will move and respond. That capability is the missing piece for machines that assist in delicate physical tasks, surgery among them, and it is where I want to take this stack next.
A robot is, in principle, an ideal surgeon: deterministic, repeatable, precise to the smallest increment of motion. What it lacks is intelligence dexterous enough for the physics of living tissue. Closing that gap means reasoning that runs from a tissue's physical behavior up to the image and back down to safe action.