Job Openings
Research Internship: Audio Toolbox and (Micro) Speech-LLMs
STRUCTURED AUDIOPROFILING FOR CONVERSATIONAL AI AGENTS
‍
Full-time | Voice & Conversational AI | Enterprise AI | Speech AI Team
Duration: 6 Months (flexible)
Location: Switzerland (Europe), on-site at AGIGO’s Zurich Office
‍
About AGIGO
AGIGO provides the enterprise-grade conversational AI infrastructure and end-to-end toolchain to design and operate high-agency, human-like AI agents that engage directly with customers over phone, email, and text, handling complete customer interactions across support, bookings, and sales. AGIGO stands out by offering true AI sovereignty through on-premises deployment and zero exposure to third-party services. Powered by AGIGO’s proprietary technology stack, the platform delivers reliable agent operations, execution assurance, ultra-low latency, seamless enterprise integration, and predictable, token-free economics.
Founded in Switzerland in February 2025 by a team of experienced AI pioneers, AGIGO is building the infrastructure for a new generation of enterprise customer interactions, combining human-like communication with the control, reliability, and economics enterprises require at scale.
‍
Your Research Mission
The moment a caller stops speaking, a voice agent has to answer a lot of questions about what it just heard. Which language and accent? Did they switch language mid-sentence, or finish their turn, or only pause to think? Did two people talk at once? Was there a number or a name in there? Is this a real voice or a synthetic one? Today each answer comes from a separate model, so every turn pays for a chain of forward passes: slow, expensive to run, and too big for the time budget of a live turn.
Core objective: Build a highly optimized W2V2-sized model (or micro SpeechLLM) that runs in real time, and returns a highly structured and detailed AudioProfile that can help to understand the spoken audio message. Accuracy on any one task is not the hard part, since high-performance open models already exist (Whisper, PyAnnote, etc). The hard part is answering all of them on every turn inside a live call: eight specialists mean eight forward passes and eight deployments to maintain. The core question of this internship: does a frozen encoder with many small heads, or a micro speech-LLM that writes the AudioProfile as JSON and picks up new tasks from instructions is sufficient? You will build and compare both approaches and eventually help us in the deployment.
‍
What the Model Has to Answer
• Turn control: the group that matters most and is usually done worst: has the caller finished or just paused (semantic endpointing), did two voices overlap, and the distinction that decides whether an agent feels rude, or that user did not mean to stop.
• Language, speaker and channel: language and accent, code-switching, how noisy is an utterance, non-silence, echo, clipping, hold music, VAD failure, etc.
• Content triggers: does this utterance contain a number, a spelled sequence, a keyword or a named entity? Digit strings are where transcription is least reliable and errors cost most, so knowing before the transcript arrives lets the stack change how it decodes and when to confirm.
• Identity, phonetics and safety: a speaker embedding per turn, so we know whether the person on the line is still the one who authenticated; phonemes; hesitations and non-speech sounds that should never reach an LLM as words; synthetic-voice detection  and watermarking.
Phase 1: Backbone and contract
A frozen W2V2-sized encoder with one shared pass. These encoders hold different information at different depths, so you probe which layer feeds which head. Then the versioned AudioProfile schema, with a confidence per field and an explicit abstain state.
Phase 2: Core heads
Language, accent and audio quality, trained from our 168k-hour, 77-language catalog using labels the annotation pipeline already produces. There is also space for synthetic data generation or bootstrapping already available data for training specific tasks.
Phase 3: The full task set
Endpointing moves past silence thresholds to a voice-activity-projection objective, since cutting people off is the worst failure here. Then overlap, backchannel versus interruption, code-switch spans, content, identity, phonemes, spoofing and watermarking.
Phase 4: Architecture study
Compare both designs on the same tasks and serving harness. Encoder-and heads gives deterministic, quantizable output in one pass; a micro speech-LLM can take on tasks it was never trained for.
Phase 5: Optimisation and serving
INT8 quantisation, ONNX Runtime with layer fusion, then batched serving, to reach the throughput target.
Phase 6: Integration
A registered, quantised model running live after each turn and offline over the catalog, where the same heads find the recordings we need instead of sampling and hoping.
‍
Key Research Challenges
Is there a capacity ceiling? As part of the internship, you will measure which tasks reinforce one another and which conflict
Does a speaker head fit next to the phonetic ones? Short turns are hard for speaker models, and a caller moving from speakerphone to handset can look like a different person.
Micro Speech-LLMs or traditional encoder-based models? The intern will benchmark multiple approaches in order to find the ”sweet spot” of performance and latency.
‍
Your Impact
The outcome of this work will replace a pipeline approach with a single highly optimized fast pass that says what is in the utterance and whether it is safe: routing to the right language and detecting code-switching, deciding whether the caller has finished, etc.
We value original thinking and encourage you to help shape and redefine the project’s direction as your research uncovers new insights. AGIGO fosters an open, collaborative environment where ideas can evolve freely. Exceptional innovation often emerges where disciplines and perspectives intersect, and we actively support creative exploration that pushes the boundaries of what Voice-AI can achieve.
‍
What You Bring
Required
• Current Master’s or PhD student (preferred), or recent graduate in Computer Science, Machine Learning, or a related degree field
• Strong Python programming skills and Git
• Solid understanding of ML fundamentals and MLOps
• Hands-on experience with PyTorch
• Expertise using Claude Code/Codex
• Fluent in English, highly motivated, willingness to learn
Bonus
• Self-supervised audio encoders and audio benchmarks
• Inference optimisation (ONNX, quantisation, pruning)
• Speech, ASR or TTS background.
‍
What You Will Gain
• Direct impact on our product: your code ships in our platform, built alongside our researchers and engineers
• Mentorship: work closely with our expert team of researchers and engineers
• Top-tier AI infrastructure: access to GPU clusters with NVIDIA Hopper (H200) and Blackwell RTX 6000 PRO NVIDIA GPUs
• Research visibility: we will actively support you in publishing your work at a top-tier conference or in a journal paper
• Disciplined and inspiring research environment: a team of sharp minds grounded in expertise, autonomy, and a shared pursuit of impactful breakthroughs
• Paid internship: market-level salary, flexible hours, unlimited coffee, drinks, fruit and snacks
• Career path: this internship may lead to a full-time permanent role in AGIGO's world-class AI R&D team
‍
How to Apply
To apply, please send your resume and a brief introduction to internships@agigo.ai with the subject line:
Research Internship – Audio Toolbox & (Micro) Speech-LLMs – [Your Full Name]
‍
By submitting your application, you agree to allow AGIGO to store and process your data for recruitment purposes. Unless otherwise requested, we may retain your data for up to one year to consider you for this or other future opportunities.
‍
Research in the Field
[1] Layer-wise Analysis of a Self-supervised Speech Representation Model, 2021. https://arxiv.org/abs/2107.04734
[2] Voice Activity Projection: Self-supervised Learning of Turn-taking Events, 2022. https://arxiv.org/abs/2205.09812
[3] Qwen2-Audio Technical Report, 2024. https://arxiv.org/abs/2407.10759
‍
AGIGO™ is a registered trademark of AGIGO AG, Switzerland.
‍
‍
