Photo of Mahyar Gohari

Mahyar Gohari

ML Enthusiast, University of Brescia

Brescia, Italy

I'm currently a postdoctoral researcher at the Department of Information Engineering (DII), University of Brescia, working on efficient computer vision methods for compressed-domain images. I previously completed my PhD in the same department, focusing on audio-visual deepfake and manipulation detection. My research spans machine learning and computer vision, with a focus on practical and efficient solutions.

Selected Work

2026IEEE ICIP 2026

Efficient Object Detection on JPEG AI Pre-Reconstruction Latents

Mahyar Gohari, Alessandro Gnutti, Fabrizio Guerrini, Nicola Adami, Riccardo Leonardi

Object detection performed directly on the latent representations of the JPEG AI learned image codec, before image reconstruction, avoiding the cost of fully decoding images.

2026Computer Vision and Image Understanding

Interpretable Detection of Singing Voice Manipulations Using Audio-Language Models

Mahyar Gohari, Davide Salvi, Paolo Bestagini, Nicola Adami

Uses audio-language models to detect manipulations in singing voice recordings and describe them in natural language. Comes with a dataset of manipulated and synthetic vocal recordings covering a range of common transformations, each annotated with a detailed description of the alterations present.

2026Data in Brief

ATDD: Multi-lingual Dataset for Auto-Tune Detection in Music Recordings

Mahyar Gohari, Paolo Bestagini, Sergio Benini, Nicola Adami

A new multilingual dataset for telling auto-tuned music apart from genuine performances, filling a gap in existing datasets. It includes tracks in English, Mandarin, and Japanese to cover a wide range of linguistic settings.

2025IEEE ICASSP 20251st place · SVDD Challenge, ISMIR 2024

Audio Features Investigation for Singing Voice Deepfake Detection

Mahyar Gohari, Davide Salvi, Paolo Bestagini, Nicola Adami

An investigation of which audio representations and features best separate real from synthetically generated singing voices. This work achieved the highest performance at the Singing Voice Deepfake Detection Challenge at ISMIR 2024.

2024IEEE WIFS 2024

Spectrogram-Based Detection of Auto-Tuned Vocals in Music Recordings

Mahyar Gohari, Paolo Bestagini, Sergio Benini, Nicola Adami

A data-driven approach that uses triplet networks to detect Auto-Tuned songs.

2024CBMI 2024

PGNN-based Approach for Robust 3D Light Direction Estimation in Outdoor Images

Marcello Zanardelli, Mahyar Gohari, Riccardo Leonardi, Sergio Benini, Nicola Adami

A physics-guided neural network (PGNN) for precise global 3D light direction estimation. The architecture integrates an illumination model, letting the network indirectly learn geometric information and improving the accuracy of the estimated light direction.

2024Data in Brief

SynthOutdoor: A Synthetic Dataset for 3D Outdoor Light Estimation

Marcello Zanardelli, Mahyar Gohari, Riccardo Leonardi, Sergio Benini, Nicola Adami

SynthOutdoor: 39,086 high-resolution images addressing data scarcity in 3D light direction estimation.

2021Turkish Journal of Computer and Mathematics Education (TURCOMAT)

Detection and Localization of Ripe Tomatoes Using Machine Vision

Mahyar Gohari

A machine-vision model for a tomato-harvesting robot that detects and locates ripe tomatoes automatically and in real time.