Keynotes

Dhruv “DJ” Jain

K1

Beyond Accuracy: The Human Side of Sound AI

 

Abstract

Sound and speech AI has made remarkable progress in recognizing, separating, describing, and generating audio. Much of this progress is measured through model-centric metrics: accuracy, error rates, localization performance, and other benchmarks that tell us how well an algorithm performs a predefined task. But the systems we ultimately build do not end at the model output. They end with a person interpreting that output, acting on it, and using it to navigate the world.

In this talk, I argue that human experience should not be treated merely as a downstream outcome of a sound AI system, but as part of the system itself. Drawing on a series of projects spanning sound recognition, accessibility, human-AI collaboration, and personalized audio, I will show how human context, goals, behavior, and interaction can reveal limitations that are invisible in benchmark evaluations—and can fundamentally change how these systems should be designed. A person may provide contextual knowledge that a model lacks, have needs that shift across situations, or value an output that is not captured by conventional notions of correctness. These examples motivate a more holistic view of sound AI in which model performance remains important, but is considered as one part of a larger human-AI system. I will discuss how this perspective can help us design technologies that are not only more accurate but also more adaptive, useful, and effective in the lives of the people they are intended to support.

 

Biography

Dhruv “DJ” Jain is an Assistant Professor of Computer Science and Engineering at the University of Michigan. His research intersects human-computer interaction, accessibility, and audio AI. He studies how sound and speech systems can be designed around human needs and contexts, making them more accessible, adaptive, and useful in the real world. Technologies pioneered in his lab have been translated into real-world systems used at scale. For example, SoundWatch, the smartwatch app for sound recognition, has been downloaded more than 4,000 times. His indoor navigation system for visually impaired users has supported more than 100,000 interactions for blind and low-vision museum visitors in India. His work has also contributed to built-in accessibility features available on all iPhones and informed real-time captioning systems at Google. DJ is a recipient of the SIGCHI Outstanding Dissertation Award, the CRA-Microsoft Trustworthy AI Fellowship, and nine best paper and honorable mention awards at top-tier computer science conferences.

For more info: https://dhruv-jain.com/




Emily Mackevicius

K2

Who Is the Listener? Privacy-Aware, Behavior-Grounded Bioacoustics in Urban Rats

 

Abstract

Acoustic machine learning is usually organized around categories chosen by people: speech, species, events, and scenes. Free-ranging animals pose a different problem. In our work on urban rats, microphones share public spaces with people, while the signals of interest are ultrasonic vocalizations whose natural functions remain largely unknown.

This setting suggests two new directions for acoustic scene analysis. First, human speech becomes sensitive incidental data. Privacy-aware monitoring requires systems that detect and separate speech in order to remove it, with evaluation centered on privacy leakage and preservation of nonhuman signals. Second, the relevant animal categories cannot be assumed from human annotations or laboratory call labels. We need to discover which acoustic distinctions matter to rats by connecting vocalizations to movement, social encounters, environmental context, and, where available, individual identity.

I will present our preliminary work synchronizing ultrasonic audio with thermal video, movement tracking, and three-dimensional reconstructions of urban environments. The scientific questions remain open. Can acoustic representations recover distinctions that predict how an animal behaves? How should benchmarks be designed when behavioral consequences provide the ground truth? Urban bioacoustics offers a challenging test case for building acoustic systems that respect the people they encounter while learning about a world organized around another species’ priorities.

 

Biography

Emily Mackevicius is a computational neuroscientist, co-founder and director of Basis Research Institute, and leader of its Collaborative Intelligent Systems group. She earned her PhD in Brain and Cognitive Sciences at MIT, where she combined neural recordings, computational modeling to study songbird learning. As a Simons Society of Fellows Junior Fellow and postdoctoral researcher at Columbia University’s Zuckerman Institute, she studied hippocampal coding of one-shot memories during natural food-caching behavior.

Her research develops new measurement and modeling approaches when existing laboratory paradigms cannot capture the behavior of interest. Her work includes publications in Cell, eLife, Nature, and Scientific Reports on episodic memory, predictive cognitive maps, and the relationship between latent preferences and group movement. At Basis, her group combines synchronized ultrasound, thermal imaging, automated tracking, three-dimensional reconstruction, and simulation to study free-ranging urban rats in New York City and Boston. She is a recipient of an NIH K99/R00 Pathway to Independence Award.

For more info: https://emackev.github.io/