Propio Language Services
Senior Machine Learning Engineer, Speech & LLM Training Data
- LocationUnited States
- TypeFull-time
- Posted2026-09-29
- Valid through2026-10-13
- BudgetUSD 11000–14500 / month
- Long-termYes
About this task
Responsibilities:
- Define the data roadmap for multilingual speech, translation, multimodal LLM and conversational AI systems.
- Build audio-processing pipelines for resampling, channel handling, voice activity detection, diarization, language identification, transcription, alignment and quality filtering.
- Curate large-scale datasets through cleaning, deduplication, PII/PHI redaction, quality scoring, sampling, balancing, versioning and lineage tracking.
- Design annotation guidelines, QA rubrics, golden datasets and reviewer workflows. Build evaluation datasets, investigate model failures and use findings to improve training data.
- Run model training, fine-tuning, post-training and evaluation experiments, including SFT, preference data, DPO/RLHF-style workflows and synthetic data generation. Productionize secure, reproducible data and ML workflows on AWS.
Requirements:
- Bachelor's or master's degree in a relevant technical field, or equivalent practical experience; at least five years in ML engineering, speech/audio ML, ML data engineering, NLP or LLM training-data workflows.
- Hands-on proficiency with Python, SQL, Linux, Git and Docker; experience with PyTorch, Hugging Face or comparable frameworks, and FFmpeg and audio-processing libraries such as TorchCodec, torchaudio or librosa.
- Experience with VAD, diarization, ASR, forced alignment, language identification and audio-quality analysis; large-scale pipelines using Databricks/Spark and Parquet/Arrow; and AWS S3, SageMaker, Glue, Step Functions, IAM and KMS.
- Experience with annotation platforms such as Labelbox, Label Studio, Scale AI, Prodigy or Argilla, and experiment-tracking or versioning tools such as MLflow, Weights & Biases, DVC, Delta Lake or LakeFS.
- Experience with multilingual speech, translation, annotation workflows and evaluation datasets. Preferred experience includes telephony or healthcare audio; Silero VAD, pyannote, WhisperX, NeMo or Kaldi; Ray or PySpark; secure handling of HIPAA-regulated or sensitive data; low-resource languages; and synthetic data, active learning, weak supervision or LLM-as-judge evaluation.
Employment type: Full-time.