Trivita AI is hiring an Intern / Junior AI Engineer (Speech AI) to research and develop ASR, TTS, speech enhancement, and voice technologies.
Job description
- Research and develop Speech AI solutions related to:
- Speech-to-Text (ASR)
- Text-to-Speech (TTS)
- Speech Enhancement / Denoising
- Voice Activity Detection (VAD)
- Speaker Diarization / Speaker Recognition
- Build pipelines for audio data processing and evaluation.
- Preprocess, label, validate, and analyze speech data.
- Experiment with open-source or commercial models such as Whisper, wav2vec 2.0, MMS, NVIDIA NeMo, CosyVoice, and similar models.
- Fine-tune and optimize models for Vietnamese and real-world use cases.
- Evaluate model quality using metrics such as WER, CER, MOS, latency, and throughput.
- Collaborate with Backend and MLOps teams to deploy models as APIs or production services.
- Read technical documentation and research papers, and stay up to date with emerging Speech AI technologies.
Intern requirements
- Third-year or final-year students, or recent graduates in Computer Science, Artificial Intelligence, Data Science, Electrical/Computer Engineering, or related fields.
- Basic knowledge of Python, Machine Learning, and Deep Learning.
- Knowledge of signal processing or audio processing is a plus.
- Experience using frameworks such as PyTorch or TensorFlow.
- Strong self-learning mindset with the initiative to experiment and read technical documentation.
Junior requirements
- 6 months to 2 years of experience in AI/ML or Speech AI.
- Proficiency in Python and at least one Deep Learning framework, preferably PyTorch.
- Understanding of one or more areas such as ASR/TTS, Audio Classification, Speech Enhancement, Voice Conversion, or Speaker Diarization.
- Experience working with audio datasets, feature extraction, and model evaluation.
- Familiarity with Git, Docker, and Linux.
- Experience deploying models as APIs is a plus.
Preferred qualifications
- Personal projects or research experience related to Speech AI.
- Experience with Whisper, wav2vec 2.0, NVIDIA NeMo, Hugging Face, or ONNX.
- Understanding of Mel-spectrograms, MFCC, FFT, sampling rates, and noise reduction.
- Experience with inference optimization techniques such as quantization, batching, streaming, or GPU serving.
- Published articles, research papers, GitHub projects, or contributions to open-source projects.
Benefits
- Work on real-world Speech AI products serving actual users.
- Receive guidance from engineers experienced in AI and production systems.
- Work with Vietnamese speech data and large-scale Speech AI problems.
- Gain hands-on experience in model training, evaluation, deployment, and performance optimization.
- Work in an environment that encourages research, experimentation, and new ideas.
- Opportunity to become a full-time employee after the internship.
Contact information
- Address: No. 01, Street 104, Quarter 3, Binh Trung Ward, Ho Chi Minh City
- Phone: 0909797699
- Email: hr@trivita.ai

