Job Openings
Research Internship: Conversational Environments and World Models for Voice AI Agents
‍AGENT-USER WORLD MODELS AND SIMULATORS
‍
Full-time | Voice & Conversational AI | Enterprise AI | Speech AI Team
Duration: 6 Months (flexible)
Location: Switzerland (Europe), on-site at AGIGO’s Zurich Office
‍
About AGIGO
AGIGO provides the enterprise-grade conversational AI infrastructure and end-to-end toolchain to design and operate high-agency, human-like AI agents that engage directly with customers over phone, email, and text, handling complete customer interactions across support, bookings, and sales. AGIGO stands out by offering true AI sovereignty through on-premises deployment and zero exposure to third-party services. Powered by AGIGO’s proprietary technology stack, the platform delivers reliable agent operations, execution assurance, ultra-low latency, seamless enterprise integration, and predictable, token-free economics.
Founded in Switzerland in February 2025 by a team of experienced AI pioneers, AGIGO is building the infrastructure for a new generation of enterprise customer interactions, combining human-like communication with the control, reliability, and economics enterprises require at scale.
‍
Your Research Mission
Real call-center recordings are the best training data for a voice agent and the hardest to get: expensive to collect, restricted by privacy rules, and almost never annotated the way we need. The calls that decide whether a deployment works are also the rare ones, such as a caller talking over the agent, a line bad enough to force a repeat, or a code-switched utterance. We cannot wait for enough of those to appear in recorded traffic, and the usual shortcut of having an LLM write a script for a TTS to read produces flat, unlabelled audio that moves no production metric. The size of the problem is measurable: voice agents complete 31–51% of the tasks their text counterparts complete at roughly 85%.
‍
What You Will Build
In this internship, you will build a configurable world model for conversational AI Agents.
Find below some requirements:
• a user simulator, a persona is a dependency graph, where each attribute is drawn conditionally and recursively;
• a client database of customers, accounts, business rules and the expected outcome. Since it knows the answer, task success and factual accuracy become computable labels. Samples might be imperfect, so we can test what an agent does when there is caller/record mismatch;
• an agent configuration: policy strictness, what it may authorize, which tools it has and how often they fail, how fresh its knowledge is, and the behaviors with no text equivalent such as speech endpointing, barge-in, etc;
• a domain lexicon: the entity pool the domain uses, which fills the world with plausible values and controls how hard they are to recognise;
• an interaction and channel layer: challenging settings such as: backchannels, interruptions, overlap, hesitation and repair, then codecs, packet loss, reverberation and noise.
Phase 1: Environment Core
A sampler that plans a call top-down and an LLM driver that writes it in a generator–critic loop, with a user simulator that keeps a goal stack so long calls stay coherent. We set the InitialConditions and WorldState schemas plus per-turn intent, dialogue-act and entity labels.
Phase 2: Personas
The dependency graph above, built so detail costs nothing in domains that do not need it and each domain sees only the attributes relevant to it. One population is then reusable everywhere.
Phase 3: Interaction and acoustics
A registry of transforms you switch on one at a time, covering floor management, pauses that break endpointing, hesitation and repair, background events and the channel, each writing a timestamped label next to the audio.
Phase 4: Agent configuration and faults
The agent-side settings above, plus a seeded fault plan for timeouts, stale responses, schema drift and authorization failures, reporting detection and recovery separately.
Phase 5: Domains, validation and the closed loop
We will work on multiple domains, each shipping asa configuration package. Then, you will aim at using synthetic data to improve real ASR, NLU and agent behavior, judged downstream, since perceptual quality does not predict how useful a corpus is for training.
‍
Key Research Challenges
Does synthetic evaluation predict real behavior? We can claim only as much as the measured correlation between synthetic curves and real calls supports.
Does anything trained on this transfer? Models trained on synthetic speech pick up synthesis artifacts, so every gain has to be checked against a strong classical baseline on real recordings.
Can a simulated caller be difficult enough? Default LLM users are too cooperative, and agents tuned against them do badly with real people. That hurts more in speech, since a cooperative caller stops producing the overlaps and hesitations we need. Also open: dual control, where the agent talks the caller through something on their own device, and environments whose state moves while the agent is still speaking.
‍
Your Impact
The output of this internship would be hard-labeled conversation data for every speech model we train, plus a set of regression harnesses. We find out how the agent handles interruptions, backchannel, noise, real semantic endpointing detection, etc.
We value original thinking and encourage you to help shape and redefine the project’s direction as your research uncovers new insights. AGIGO fosters an open, collaborative environment where ideas can evolve freely. Exceptional innovation often emerges where disciplines and perspectives intersect, and we actively support creative exploration that pushes the boundaries of what Voice-AI can achieve.
‍
What You Bring
Required
- Current Master’s or PhD student (preferred), or recent graduate in Computer Science, Machine Learning, or a related degree field
- Strong Python programming skills and Git
- Solid understanding of ML fundamentals and MLOps
- Hands-on experience with PyTorch
- Expertise using Claude Code/Codex
- Fluent in English, highly motivated, willingness to learn
Bonus
- Experience with Hugging Face models (for LLMs, ASR, or "speech-LLMs")
- Hands-on experience with audio AI (ASR/TTS) model training and development
- Hands-on experience with large-scale data processing pipelines
‍
What You Will Gain
- Direct impact on our product: your code ships in our platform, built alongside our researchers and engineers
- Mentorship: work closely with our expert team of researchers and engineers
- Top-tier AIÂ infrastructure: access to GPU clusters with NVIDIA Hopper (H200) and Blackwell RTX 6000 PRO NVIDIA GPUs
- Research visibility: we will actively support you in publishing your work at a top-tier conference or in a journal paper
- Disciplined and inspiring research environment: a team of sharp minds grounded in expertise, autonomy, and a shared pursuit of impactful breakthroughs
- Paid internship: market-level salary, flexible hours, unlimited coffee, drinks, fruit and snacks
- Career path: this internship may lead to a full-time permanent role in AGIGO's world-class AIÂ R&D team
‍
‍How to Apply
To apply, please send your resume and a brief introduction to internships@agigo.ai with the subject line:
‍Research Internship – Conversational Environments & World Models for Voice AI-Agents – [Your Full Name]
‍
By submitting your application, you agree to allow AGIGO to store and process your data for recruitment purposes. Unless otherwise requested, we may retain your data for up to one year to consider you for this or other future opportunities.
Research in the Field
[1] Ď„-Voice: Benchmarking Full-Duplex Voice Agents on Real-World Domains, 2026. https://arxiv.org/abs/2603.13686
[2] Beyond Cooperative Simulators: Generating Realistic User Personas forRobust Evaluation of LLM Agents, 2026. https://arxiv.org/abs/2605.12894
[3] Gaia2: LLM Agents on Dynamic and Asynchronous Environments, ICLR 2026. https://arxiv.org/abs/2602.11964
‍
AGIGO™ is a registered trademark of AGIGO AG, Switzerland.‍
‍
