Job Openings
Research Internship: Efficient Inference for Text-to-Speech Models and LLMs
QUANTIZATION, PRUNING, SPARSITY
‍
Full-time | Voice & Conversational AI | Enterprise AI | Speech AI Team
Duration: 6 Months (flexible)
Location: Switzerland (Europe), on-site at AGIGO’s Zurich Office
‍
About AGIGO
AGIGO provides the enterprise-grade conversational AI infrastructure and end-to-end toolchain to design and operate high-agency, human-like AI agents that engage directly with customers over phone, email, and text, handling complete customer interactions across support, bookings, and sales. AGIGO stands out by offering true AI sovereignty through on-premises deployment and zero exposure to third-party services. Powered by AGIGO’s proprietary technology stack, the platform delivers reliable agent operations, execution assurance, ultra-low latency, seamless enterprise integration, and predictable, token-free economics.
Founded in Switzerland in February 2025 by a team of experienced AI pioneers, AGIGO is building the infrastructure for a new generation of enterprise customer interactions, combining human-like communication with the control, reliability, and economics enterprises require at scale.
‍
Your Research Mission
Real-time voice agents require very low-latency decoding and increased serving costs compared to offline models. The latency becomes an extremely differentiating aspect, since a reply from an Voice Agent arriving after one second starts feeling broken. In this internship, you will build a complete pipeline for quantizaiton/quantization-aware training, and pruning for our internal LLMs and TTS models. Ideally, one initial checkpoint is compiled into a family of variants, each valid for a particular GPU generation (ADA/Hopper/Blackwell), precision, kernel stack, and batching regime, and every variant has to clear the same speech-aware quality gate before it is allowed out. Therefore, an important question arises: given the hardware and the latency requirements, which build are we allowed to serve? This matters because our stack is not uniform, but rather fluid, with Hopper, Blackwell or ADA machines requested on demand, which also might reward different recipes, so the same model has a different best answer depending on the initial conditions.
‍
What You Will Build
• The build matrix. An initial checkpoint in, let’s say FP16/BF16 precision, then build a matrix from: weight-only 4-bit, activation quantisation with outlier handling, FP8 and NVFP4, KV-cache quantisation, and 2:4 sparsity where the hardware can use it. Each build is tagged with the hardware and workload it is valid for.
• The quality gate. You will implement strong evaluation pipelines to signal issues in a quantized checkpoint, e.g., UTMOS, word error rate or speaker similarity for TTS, or other metrics such as intent accuracy or NER performance for LLMs.
• Elastic serving and speculative decoding. One checkpoint offering several operating points chosen per call rather than per deployment; and speculative decoding for specific LLMs tasks or TTS.
Phase 1: Harness and gate
The benchmark harness and the CI quality gate, plus the schema for a build record: what it is valid for, and what it measured.
Phase 2: Post-training quantisation
With and without in-domain calibration data.
Phase 3: Recovery
Quantisation-aware training and distillation, aiming to beat post-training quantisation at the same bit-width.
Phase 4: Sparsity
2:4 pruning, and investigate whether structured sparsity becomes real throughput.
‍
Key Research Challenges
Does the best build actually differ by hardware, and by how much? If both generations rank the variants the same way.
Is speculative decoding lossless for audio? You will investigate in which conditions speculative decoding for audio is lossless.
‍
Your Impact
The resulting recipe of this internship will translate on optimizations in the models deployed in our stack, where each latency point and increase in tok/s really matters.
We value original thinking and encourage you to help shape and redefine the project’s direction as your research uncovers new insights. AGIGO fosters an open, collaborative environment where ideas can evolve freely. Exceptional innovation often emerges where disciplines and perspectives intersect, and we actively support creative exploration that pushes the boundaries of what Voice-AI can achieve.
‍
What You Bring
Required
• Current Master’s or PhD student (preferred), or recent graduate in Computer Science, Machine Learning, or a related degree field
• Strong Python programming skills and Git
• Solid understanding of ML fundamentals and MLOps
• Hands-on experience with PyTorch
• Expertise using Claude Code/Codex
• Fluent in English, highly motivated, willingness to learn
Bonus
• CUDA, mixed precision, profiling, inference servers, or quantization tools such as: llm-compressor (vLLM) or Model-Optimizer (NVIDIA).
‍
What You Will Gain
• Direct impact on our product: your code ships in our platform, built alongside our researchers and engineers
• Mentorship: work closely with our expert team of researchers and engineers
• Top-tier AI infrastructure: access to GPU clusters with NVIDIA Hopper (H200) and Blackwell RTX 6000 PRO NVIDIA GPUs
• Research visibility: we will actively support you in publishing your work at a top-tier conference or in a journal paper
• Disciplined and inspiring research environment: a team of sharp minds grounded in expertise, autonomy, and a shared pursuit of impactful breakthroughs
• Paid internship: market-level salary, flexible hours, unlimited coffee, drinks, fruit and snacks
• Career path: this internship may lead to a full-time permanent role in AGIGO's world-class AI R&D team
‍
How to Apply
To apply, please send your resume and a brief introduction to internships@agigo.ai with the subject line:
Research Internship – Efficient Inference for Text-to-Speech Models and LLMs – [Your Full Name]
‍
By submitting your application, you agree to allow AGIGO to store and process your data for recruitment purposes. Unless otherwise requested, we may retain your data for up to one year to consider you for this or other future opportunities.
‍
Research in the Field
[1] AWQ: Activation-aware Weight Quantization for LLM Compression, MLSys 2024. https://arxiv.org/abs/2306.00978
[2] SmoothQuant: Post-Training Quantization for Large Language Models, 2022. https://arxiv.org/abs/2211.10438
[3] Principled Coarse-Grained Acceptance for Speculative Decoding in Speech, ICASSP 2026. https://arxiv.org/abs/2511.13732
‍
AGIGO™ is a registered trademark of AGIGO AG, Switzerland.
‍
‍
