
Lester Phillip Violeta
Research Scientist (Speech LLMs & Post-Training)
DubGuild, Tokyo, Japan
I am a Research Scientist at DubGuild specializing in speech LLMs. I worked on the training and evaluation of conversational speech models for synthetic dialogue generation, through full-parameter supervised fine-tuning and reinforcement-learning post-training (YouTube demo). I also helped scale an English-Japanese speech LLM (check our blogs).
I received my Ph.D. in Informatics from Nagoya University at Toda Laboratory under the supervision of Professor Tomoki Toda. My research on speech synthesis, voice conversion, and speech recognition has been published at ICASSP, Interspeech, EUSIPCO, ASRU, and in IEEE journals. I was the head organizer of the Singing Voice Conversion Challenge 2025 and a member of the organizing committee for the 2023 challenge, and I serve on the peer-review committees of conferences such as ASRU, SLT, ICASSP, Interspeech, and IJCNN, and journals like IEEE JSTSP.
I have a deep international background now based in Japan, having done my B.S. in the Philippines and a research exchange in France. Outside of research, I like bouldering (check out my instagram page) and learning Japanese.
News
🎓 Graduated with my Ph.D. from Nagoya University!
Experience
Research Scientist — DubGuild
Focused on the training and evaluation of conversational speech LLMs for synthetic dialogue generation, including full-parameter SFT and RL post-training.
Research Engineer — CoeFont
Developed real-time voice conversion models for the CoeFont Voice Changer and trained large-scale emotional TTS models.
ML Engineer — Voice-Swap.AI
Developed singing voice conversion models for the Voice-Swap singing studio, used by music-industry clients.
Research Assistant — Sony CSL Tokyo
Manager: Dr. Taketo Akama
Researched highly controllable, low-resource singing voice synthesis.
Research Intern — NTT Media Intelligence Laboratories
Manager: Dr. Atsushi Ando
Developed and analyzed speaker diarization systems using various encoders.
Research Intern — Hitachi Ltd.
Manager: Dr. Takashi Sumiyoshi
Developed speech recognition systems for low-resource datasets.
Technical Focus
ML Frameworks
PyTorch, Megatron-Bridge, NeMo AutoModel, verl, TRL, vLLM
Research Areas
SFT, RL post-training (GRPO, DPO, KTO), reward design
Infrastructure
Multi-node GPU training (FSDP, expert parallelism, context parallelism)
Languages
English (native), Tagalog (native), Japanese (upper intermediate)
Publications

Technical Report 2026
Scaling Japanese Speech Foundation Models and Examining TTS Performance (in Japanese)
長谷川 直哉, 相田 優希, 廣岡 聖司, 林 春太朗, Lester Phillip Violeta, 大嶽 匡俊

ICASSP 2026
The Singing Voice Conversion Challenge 2025: From Singer Identity Conversion To Singing Style Conversion
Lester Phillip Violeta, Xueyao Zhang, Jiatong Shi, Yusuke Yasuda, Wen-Chin Huang, Zhizheng Wu, Tomoki Toda

IEEE JSTSP 2025
Resolving Domain Mismatches in Electrolaryngeal Speech Enhancement With Linguistic Intermediates
Lester Phillip Violeta, Wen-Chin Huang, Ding Ma, Ryuichi Yamamoto, Kazuhiro Kobayashi, Tomoki Toda

ICASSP 2024
Electrolaryngeal Speech Intelligibility Enhancement through Robust Linguistic Encoders
Lester Phillip Violeta, Wen-Chin Huang, Ding Ma, Ryuichi Yamamoto, Kazuhiro Kobayashi, Tomoki Toda

IEEE/ACM TASLP 2024
Pretraining and Adaptation Techniques for Electrolaryngeal Speech Recognition
Lester Phillip Violeta, Ding Ma, Wen-Chin Huang, Tomoki Toda

Technical Report 2024
A Preliminary Investigation on Flexible Singing Voice Synthesis Through Decomposed Framework with Inferrable Features
Lester Phillip Violeta, Taketo Akama

ASRU 2023
The Singing Voice Conversion Challenge 2023
Wen-Chin Huang, Lester Phillip Violeta, Songxiang Liu, Jiatong Shi, Tomoki Toda

APSIPA 2023
An Analysis of Personalized Speech Recognition System Development for the Deaf and Hard-of-hearing
Lester Phillip Violeta, Tomoki Toda

ICASSP 2023
Intermediate Fine-tuning Using Imperfect Synthetic Speech for Improving Electrolaryngeal Speech Recognition
Lester Phillip Violeta, Ding Ma, Wen-Chin Huang, Tomoki Toda

Interspeech 2022
Investigating Self-Supervised Pretraining Frameworks for Pathological Speech Recognition
Lester Phillip Violeta, Wen-Chin Huang, Tomoki Toda
Education
Nagoya University, Japan
Ph.D. in Informatics
Advisor: Prof. Tomoki Toda
Thesis: Domain Adaptation Techniques for Electrolaryngeal Speech Recognition and Enhancement
Nagoya University, Japan
M.S. in Informatics
Advisor: Prof. Tomoki Toda
Thesis: Pretraining and Adaptation Techniques for Pathological Speech Recognition
Ateneo de Manila University, Philippines
B.S. Electronics Engineering
Institut catholique d'arts et métiers — Site de Paris-Sénart, France
Research Exchange Semester
Academic Service & Awards
Head Organizer
The Singing Voice Conversion Challenge 2025
ICASSP 2026
Organizing Committee
The Singing Voice Conversion Challenge 2023
ASRU 2023 Special Session
Peer Review Committee
IEEE SLT, IEEE ICASSP, ISCA Interspeech, IEEE IJCNN, IEEE ASRU, IEEE JSTSP
Monbukagakusho Japanese Government Scholarship (Ph.D.)
Monbukagakusho Japanese Government Scholarship (Master’s)
Interspeech 2022 Travel Grant
Nagoya University Interdisciplinary Frontier Fellowship
