Publications
Generated automatically from DBLP, use Google Scholar for the most up-to-date information.
2026
-
Anamnesis: An Open-Source Platform for Large-Scale Backstory-Conditioned Survey Simulation
arXiv 2026 -
Animation2Code: Evaluating Temporal Visual Reasoning in Video-to-Code Generation
arXiv 2026 -
Attend Before Attention: Efficient and Scalable Video Understanding via Autoregressive Gazing
arXiv 2026 -
Auditing Protocol-Level Shortcuts in Large Audio Language Model Judges for Speech Evaluation
arXiv 2026 -
DigitalCoach: Communication and Grounding Gaps in Human and Agentic Computer Use Coaching
arXiv 2026 -
PitchBench: Measuring Pitch Hearing in Audio-Language Models
arXiv 2026 -
Playful Agentic Robot Learning
arXiv 2026 -
ScribbleEdit: Synthetic Data for Image Editing with Scribbles and Text
arXiv 2026 -
Stateful Visual Encoders for Vision-Language Models
arXiv 2026 -
T-Rex: Tactile-Reactive Dexterous Manipulation
arXiv 2026 -
Unintended Effects of Geographic Conditioning in Large Language Models
arXiv 2026 -
VisGym: Diverse, Customizable, Scalable Environments for Multimodal Agents
arXiv 2026
2025
-
Do Statistical Patterns in Neural Audio Codec Tokens by Synthesized Speech Reveal Structure beyond Speech Quality?
JASA 2025 DOI -
A Roadmap for Alignable Algorithmic Decision-Makers in the Medical Triage Domain
IEEE CAI 2025 -
Analysing the Language of Neural Audio Codecs
-
CLAIRA: Leveraging Large Language Models to Judge Audio Captions
-
Do What? Teaching Vision-Language-Action Models to Reject the Impossible
Findings of EMNLP 2025 arXiv ACL Anthology -
Enough Coin Flips Can Make LLMs Act Bayesian
ACL 2025 arXiv ACL Anthology -
Generate, but Verify: Reducing Hallucination in Vision-Language Models with Retrospective Resampling
-
LISAt: Language-Instructed Segmentation Assistant for Satellite Imagery
-
Puzzled by Puzzles: When Vision-Language Models Can't Take a Hint
-
REOrdering Patches Improves Vision Models
-
Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark
ICLR 2025 OpenReview -
TULIP: Contrastive Image-Text Learning with Richer Vision Understanding
ICCV Workshops 2025 PDF -
Are Large Reasoning Models Interruptible?
arXiv 2025 -
Constantly Improving Image Models Need Constantly Improving Benchmarks
arXiv 2025 -
From Talk to Triage: Pluralism is Necessary but Not Sufficient for AI Alignment
OSF Preprints 2025 DOI -
Higher-Order Binding of Language Model Virtual Personas: a Study on Approximating Political Partisan Misperceptions
arXiv 2025
2024
-
ALOHa: A New Measure for Hallucination in Captioning Models
-
An Efficient Self-Learning Framework For Interactive Spoken Dialog Systems
ICML 2024 arXiv PMLR OpenReview -
Anim-400K: A Large-Scale Dataset for Automated End to End Dubbing of Video
-
Distribution Aware Metrics for Conditional Natural Language Generation
LREC-COLING 2024 arXiv ACL Anthology -
Multi-Stage Multi-Modal Pre-Training for Automatic Speech Recognition
LREC-COLING 2024 ACL Anthology -
See, Say, and Segment: Teaching LMMs to Overcome False Premises
-
Task Oriented Dialogue as a Catalyst for Self-Supervised Automatic Speech Recognition
-
Virtual Personas for Language Models via an Anthology of Backstories
EMNLP 2024 arXiv ACL Anthology -
Automatic Audio Captioning with Encoder Fusion, Multi-Layer Aggregation, and Large Language Model Enriched Summarization
DCASE 2024 Judges' Award -
Multimodal Attention Merging for Improved Speech Recognition and Audio Event Classification
-
Analyzing The Language of Visual Tokens
arXiv 2024 -
Rediscovering the Latent Dimensions of Personality with Large Language Models as Trait Descriptors
arXiv 2024 -
Understanding, Building, and Evaluating Models for Context Aware Conditional Natural Language Generation
PhD thesis, UC Berkeley 2024
2023
-
CLAIR: Evaluating Image Captions with Large Language Models
EMNLP 2023 arXiv ACL Anthology -
Domain Adaptation with External Off-Policy Acoustic Catalogs for Scalable Contextual End-to-End Automated Speech Recognition
ICASSP 2023 PDF -
IC3: Image Captioning by Committee Consensus
-
Towards Understanding How Machines Can Learn Causal Overhypotheses
CogSci 2023 arXiv eScholarship -
Scalable and Accurate Self-Supervised Multimodal Representation Learning without Aligned Video and Text Data
WACV Workshops 2023
2022
-
Content-Context Factorized Representations for Automated Speech Recognition
-
Learning Causal Overhypotheses through Exploration in Children and Computational Models
-
Multi-Modal Pre-Training for Automated Speech Recognition
-
Multimodal Semantic Mismatch Detection in Social Media Posts
MMSP 2022 -
What's in a Caption? Dataset-Specific Linguistic Diversity and Its Effect on Visual Description Models and Metrics
-
LAVA: Language Audio Vision Alignment for Contrastive Video Pre-Training
arXiv 2022 -
Misinformation Detection in Social Media Video Posts
arXiv 2022 -
Hallucination Is All You Need: Using Generative Models for Test Time Data Augmentation
Technical report 2022
2021
-
Conversational Physical Activity Coaches for Spanish and English Speaking Women: A User Design Study
Frontiers in Digital Health 2021 DOI -
A Semantic Segmentation Network for Urban-Scale Building Footprint Extraction Using RGB Satellite Imagery
arXiv 2021 -
Exploring the Effects of View Transforms on Self-Supervised Video Representation Learning Techniques
Technical report 2021
2020
2019
2018
-
Rapid Randomized Restarts for Multi-Agent Path Finding Solvers
SOCS 2018 PDF -
T-SNE-CUDA: GPU-Accelerated T-SNE and its Applications to Modern Data
-
Diagnostic Visualization for Deep Neural Networks Using Stochastic Gradient Langevin Dynamics
arXiv 2018 -
Leveraging Class Similarity to Improve Deep Neural Network Robustness
arXiv 2018
2017
2016
-
Going Deeper in Facial Expression Recognition Using Deep Neural Networks
WACV 2016 -
Facial Expression Recognition from World Wild Web
CVPR Workshops 2016