ICASSP: What “signal processing” has come to mean

Alexa scientist Ariya Rastrow on the blurring boundaries between acoustic processing and language understanding.

The International Conference on Acoustics, Speech, and Signal Processing (ICASSP), which starts today, is now in its 45th year, and according to Google Scholar’s rankings, it’s the highest-impact conference in the field of signal processing.

But as speech-related technologies have matured, the definition of signal processing has expanded. “ICASSP is a mix of a lot of different tracks,” says Ariya Rastrow, an Alexa principal research scientist who attended his first ICASSP in 2006. “It has the whole spectrum, from very low-level signal processing all the way to interpretation and natural-language understanding.”

Ariya Rastrow.png
Alexa senior principal scientist Ariya Rastrow
Credit: Jordan Stead

This diversity, Rastrow explains, simply reflects that of the human audio-processing system. The brain doesn’t rely exclusively on acoustic signals to recognize words, and neither should computer systems.

“The interaction between language and acoustics is very dynamic from the human perspective,” Rastrow says. “If I’m talking to you in a very clean environment, we are capable of following on the acoustic level at very high resolution. But if we’re sitting in a noisy bar, you as a human are going to rely more on your prior — on a semantic level, what are the things that the other person might say? what are the topics that they might talk about? —and use that to enhance your recognition.”

Traditionally, the task of spoken-language understanding has been broken into two components: automatic speech recognition (ASR), which converts an acoustic speech signal into text, and natural-language understanding (NLU), which makes sense of the text.

But in fact, speech recognition usually relies on higher-level linguistic features to identify words. The traditional ASR system consists of an acoustic model, which translates acoustic signals into low-level phonetic representations; a lexicon, which maps sequences of low-level phonetic representations to words; and a language model, which uses high-level statistics about words’ co-occurrence to adjudicate between competing interpretations of the acoustic signal.

“Twenty, twenty-five years ago, there was this pragmatic idea to build factored systems,” Rastrow explains. “You have clear-cut boundaries between components of the system. Traditional speech recognition systems are built over an architecture that we call a hidden Markov model (HMM) architecture. The HMM architecture will put these multiple knowledge sources together at inference time. But the acoustic model and the language model are trained separately.”

Shared representations

Recently, however, this approach has begun to give way to end-to-end training of large, neural-network-based architectures. That is, a single neural network is trained on examples that consist of acoustic inputs and fully transcribed outputs, and it directly learns the relationships previously encoded in the ASR system’s separate components.

“This has many benefits” Rastrow says, “one being that by doing joint training you build systems that are more optimized in terms of accuracy. If you build factored systems, often you train each component for a specific objective function, and at inference time, they don’t know how to handle disfluencies and errors. By virtue of advances in architectures and doing joint training and multitask training, the systems are becoming more robust to those types of confusions.”

“That’s one benefit,” Rastrow continues. “Another is that the system gains in efficiency. By having a mechanism to do knowledge transfer, joint training, or shared representation, you get to the point where different parts of the systems can rely on the same types of representations or shared layers [of the network]. This can result in compression of the overall size of the system, execution speedups, and opportunities to deploy such systems on low-resource devices and hardware.

“For example, if you’re doing acoustic-event detection, and you’re also doing wake word detection and whisper detection, which are different types of audio-based classification tasks, one way is to build all the systems separately. The other way is that you can do knowledge transfer and shared representation learning, and by virtue of those shared network components and layers, you can gain efficiency beyond the obvious accuracy improvements.

“Also, the whole system is done in neural-network execution that we know how to accelerate both on the software and the hardware side, versus this explicit knowledge representation — lexicon versus language model. Traditionally, these are not deep-learning based, so we could not leverage these efficiency mechanisms. For the last two to three years, we have been pursuing this direction.”

Total integration

Allowing a single large model to integrate the ASR system’s low-level acoustic-signal processing and high-level language modeling raises the prospect of taking advantage of still higher-level linguistic features. In one of the 19 Amazon papers at this year’s ICASSP, for instance, Alexa researchers report using semantic features to help distinguish between utterances intended for Alexa and those that are not, where in the past, Alexa’s “device directedness” detector relied solely on acoustic features.

The end point of all this integration, of course, would be a single neural network that executed the entire task of spoken-language understanding — both ASR and NLU.

“There is emerging research that shows that at least for a subset of interactions, you can build a single, small-footprint network that can directly translate audio to the semantic level,” Rastrow says. “You get even better latency. You don’t have to do stage-wise execution. Also, there are studies showing that humans don’t do recognition word by word. We carry information on the parts of the speech that are semantically important for the topic, for the conversation.”

“But challenges remain,” Rastrow says. “These all-neural systems thrive on data. And once you move closer to the understanding layer, you have to cope more and more with data sparsity and the nuances of unique interactions. On the acoustic level, for the sound <p>, even across languages, you can get a lot of examples. But as you go closer to the semantic and sentence-level understanding, the patterns become more unique.

“One challenge is how we combine these new architectures for doing direct audio to NLU with our advances in semi-supervised learning and unsupervised learning. Another challenge is how to combine very data-oriented learning systems with some kind of reasoning or logic.

“I’ll give you an example. If you say, ‘Alexa turn on the bedroom light’, and Alexa misinterprets and turns on the kitchen light, and you follow that by saying, ‘No, Alexa, don’t turn on the kitchen light,’ now you have the negation problem. When you say ‘Don’t turn it on’, you really mean ‘Turn it off’. It is very hard to find those examples in data. Traditionally, we know how to address that problem with rules and logic and reasoning, but relying merely on data might not give us a good representation of those unique patterns. So the questions in the next two, three years of research will be how to combine those systems with either semi-supervised or unsupervised learning and how to combine them with knowledge and logic.”

Research areas

Related content

US, NY, New York
We are seeking an Applied Scientist to lead the development of evaluation frameworks and data collection protocols for robotic capabilities. In this role, you will focus on designing how we measure, stress-test, and improve robot behavior across a wide range of real-world tasks. Your work will play a critical role in shaping how policies are validated and how high-quality datasets are generated to accelerate system performance. You will operate at the intersection of robotics, machine learning, and human-in-the-loop systems, building the infrastructure and methodologies that connect teleoperation, evaluation, and learning. This includes developing evaluation policies, defining task structures, and contributing to operator-facing interfaces that enable scalable and reliable data collection. The ideal candidate is highly experimental, systems-oriented, and comfortable working across software, robotics, and data pipelines, with a strong focus on turning ambiguous capability goals into measurable and actionable evaluation systems. Key job responsibilities - Design and implement evaluation frameworks to measure robot capabilities across structured tasks, edge cases, and real-world scenarios - Develop task definitions, success criteria, and benchmarking methodologies that enable consistent and reproducible evaluation of policies - Create and refine data collection protocols that generate high-quality, task-relevant datasets aligned with model development needs - Build and iterate on teleoperation workflows and operator interfaces to support efficient, reliable, and scalable data collection - Analyze evaluation results and collected data to identify performance gaps, failure modes, and opportunities for targeted data collection - Collaborate with engineering teams to integrate evaluation tooling, logging systems, and data pipelines into the broader robotics stack - Stay current with advances in robotics, evaluation methodologies, and human-in-the-loop learning to continuously improve internal approaches - Lead technical projects from conception through production deployment - Mentor junior scientists and engineers
US, WA, Seattle
Prime Video is a first-stop entertainment destination offering customers a vast collection of premium programming in one app available across thousands of devices. Prime members can customize their viewing experience and find their favorite movies, series, documentaries, and live sports – including Amazon MGM Studios-produced series and movies; licensed fan favorites; and programming from Prime Video subscriptions such as Apple TV+, HBO Max, Peacock, Crunchyroll and MGM+. All customers, regardless of whether they have a Prime membership or not, can rent or buy titles via the Prime Video Store, and can enjoy even more content for free with ads. Are you interested in shaping the future of entertainment? Prime Video's technology teams are creating best-in-class digital video experience. As a Prime Video team member, you’ll have end-to-end ownership of the product, user experience, design, and technology required to deliver state-of-the-art experiences for our customers. You’ll get to work on projects that are fast-paced, challenging, and varied. You’ll also be able to experiment with new possibilities, take risks, and collaborate with remarkable people. We’ll look for you to bring your diverse perspectives, ideas, and skill-sets to make Prime Video even better for our customers. With global opportunities for talented technologists, you can decide where a career Prime Video Tech takes you! As an Applied Scientist, you will apply state of the art natural language processing and computer vision research to video centric digital media. We are looking for scientists with expertise in vision-language models/multimodal LLMs and long-form content understanding (full movies/episode vs. short clips). You will be dealing with architectures that handle long-context understanding and causal reasoning across extended temporal sequences. Key job responsibilities Our team builds multi-modal machine learning technologies to enrich and understand video content. We aim not only to understand individual components within the content itself, but also their relationships to each other to provide a holistic and broader contextual understanding. This powers the next generation of video understanding and search capabilities for Prime Video. About the team Prime Video's Content Localization, Understanding & Enrichment organization is responsible for 1) enabling Prime Video to "see" and "understand" video content including characters, scenes, dialogue, events & visual elements and 2) delivering localized, accessible content that meets a consistent cinematic quality standard at scale. This team's mission is to deeply understand all content and empower all customers with relevant language options, innovative accessibility assists, and rich title-information across all their content-experiences on Prime Video. We create and publish content on-time that's meaningful, accurate, and accessible to every customer globally. We delight our customers by pushing the boundaries of content understanding and enrichment. Through inclusion and innovation, we do the most fulfilling work of our career.
IN, KA, Bengaluru
Do you want to join an innovative team of scientists who use machine learning and statistical techniques to create state-of-the-art solutions for providing better value to Amazon’s customers? Do you want to build and deploy advanced algorithmic systems that help optimize millions of transactions every day? Are you excited by the prospect of analyzing and modeling terabytes of data to solve real world problems? Do you like to own end-to-end business problems/metrics and directly impact the profitability of the company? Do you like to innovate and simplify? If yes, then you may be a great fit to join the Machine Learning and Data Sciences team for India Consumer Businesses. If you have an entrepreneurial spirit, know how to deliver, love to work with data, are deeply technical, highly innovative and long for the opportunity to build solutions to challenging problems that directly impact the company's bottom-line, we want to talk to you. Major responsibilities - Use machine learning and analytical techniques to create scalable solutions for business problems - Analyze and extract relevant information from large amounts of Amazon’s historical business data to help automate and optimize key processes - Design, development, evaluate and deploy innovative and highly scalable models for predictive learning - Research and implement novel machine learning and statistical approaches - Work closely with software engineering teams to drive real-time model implementations and new feature creations - Work closely with business owners and operations staff to optimize various business operations - Establish scalable, efficient, automated processes for large scale data analyses, model development, model validation and model implementation - Mentor other scientists and engineers in the use of ML techniques
US, WA, Seattle
Amazon Seller Assistant is our flagship GenAI-first, multi-agent system that reimagines Seller experience. Our vision is to provide each seller with a proactive, autonomous, agentic assistant that understands their business and helps them navigate the complexities of selling by anticipating their needs, surfacing insights, resolving issues, taking actions on their behalf, and helping them grow. Amazon Seller Assistant helps millions of sellers on Amazon serve billions of customers worldwide. We are seeking a world-class Senior Data Scientist to help define and build the next generation of Amazon Seller Assistant. You will partner with top-tier scientist, engineers and product teams to launch production-grade agentic capabilities at Amazon's scale — owning your problem space end-to-end, from a crisp customer insight to a shipped product that millions of sellers rely on. Key job responsibilities • Own the science vision, strategy, and roadmap for a key Seller Assistant capability area. • Define and ship agentic experiences — sub-agent onboarding, tool onboarding, evaluations— that solve hard seller problems at scale. • Partner with scientists and engineers to translate frontier AI research into production-grade features sellers trust and depend on. • Design rigorous evaluation frameworks — automated and human-in-the-loop — to measure agent quality, accuracy, and business impact. • Deep-dive into seller data, identify unmet needs, and write compelling PRFAQs that set the direction for your team. • Drive cross-functional alignment across science, engineering, UX, and business teams to deliver with speed and quality. About the team Amazon Seller Assistant team operates at the very frontier of agentic AI and agentic commerce — not as a research group, but as a team shipping production-grade, multi-agent systems used by millions of sellers worldwide. We move with the urgency of a startup and the resources of the world's most customer-obsessed company, the latest breakthroughs in science and engineering into capabilities that sellers rely on every day.
US, NY, New York
MULTIPLE POSITIONS AVAILABLE Employer: Amazon Development Center U.S., Inc. Offered Position: Applied Scientist III - AMZ007408 Job Location: New York, NY Position Responsibilities: Participate in the design, development, evaluation, deployment, and updating of formal reasoning systems for security, privacy, and data protection applications. Drive technical and scientific innovation in security automation, data protection, and privacy-preserving technologies, with a focus on developing scalable solutions for cloud environments. Develop and/or apply formal verification techniques and automated theorem proving methods for different applications in cloud security and privacy. Collaborate with internal and external users to understand requirements and enhance formal verification and automated reasoning capabilities. Lead research and development efforts in AI security, specifically evaluate emerging threats and opportunities, including securing Generative AI systems and designing robust safeguards. Proactively identify and explore new opportunities for deploying and leveraging formal reasoning solutions across various domains.
US, CA, San Francisco
The Amazon Center for Quantum Computing (CQC) is seeking to hire an Applied Science Manager to lead a team of scientists in the physical design and simulation of superconducting quantum processors. In this role, you will use advanced modeling, simulation, and experimental design to drive improvements in scaling and performance. You will partner with other physics and engineering teams to advance the development of fault-tolerant quantum computers. Key job responsibilities - Hire Applied Scientists from diverse technical backgrounds to design quantum processors and improve the design process - Develop scientific talent through goal setting, feedback, collaborative work, and coaching - Collaborate with other science teams in designing experiments to overcome scaling and performance limitations - Influence engineering team development priorities in enabling systematic processor design and simulation workflows - Manage tactical and strategic initiatives with scientific projects pursued within team - Enable creative and innovative experimentation while striving for operational excellence About the team The Amazon Center for Quantum Computing (CQC) is a multi-disciplinary team of scientists, engineers, and technicians, on a mission to develop a fault-tolerant quantum computer. Inclusive Team Culture Here at Amazon, it’s in our nature to learn and be curious. Our employee-led affinity groups foster a culture of inclusion that empower us to be proud of our differences. Ongoing events and learning experiences, including our Conversations on Race and Ethnicity (CORE) and AmazeCon conferences, inspire us to never stop embracing our uniqueness. Diverse Experiences Amazon values diverse experiences. Even if you do not meet all of the preferred qualifications and skills listed in the job description, we encourage candidates to apply. If your career is just starting, hasn’t followed a traditional path, or includes alternative experiences, don’t let it stop you from applying. Mentorship & Career Growth We’re continuously raising our performance bar as we strive to become Earth’s Best Employer. That’s why you’ll find endless knowledge-sharing, mentorship and other career-advancing resources here to help you develop into a better-rounded professional. Work/Life Balance We value work-life harmony. Achieving success at work should never come at the expense of sacrifices at home, which is why we strive for flexibility as part of our working culture. When we feel supported in the workplace and at home, there’s nothing we can’t achieve in the cloud. Export Control Requirement Due to applicable export control laws and regulations, candidates must be either a U.S. citizen or national, U.S. permanent resident (i.e., current Green Card holder), or lawfully admitted into the U.S. as a refugee or granted asylum, or be able to obtain a US export license. If you are unsure if you meet these requirements, please apply and Amazon will review your application for eligibility.
GB, London
The Agentic Automated Reasoning Group is building the next generation of software verification tools combining advances in artificial intelligence, the computational capacity of the cloud, and our deep expertise in the domain. Join us if you want to be a part of this transformational endeavor. The Strata team (https://github.com/strata-org) is seeking an applied scientist with broad interest and expertise in model checking, interactive theorem proving, programming language semantics, and generative AI. You will combine your expertise with that of your coworkers to build new tools that solve code analysis problems previously considered beyond reach. Our application areas span all the way from Infrastructure as Code to high-performance cryptography written in assembly code, while our methods span from interactive theorem proving to automated test generation. Each day, hundreds of thousands of developers make billions of transactions worldwide on AWS. They harness the power of the cloud to enable innovative applications, websites, and businesses. Using automated reasoning technology and mathematical proofs, AWS allows customers to answer questions about security, availability, durability, and functional correctness. We call this provable security, absolute assurance in security of the cloud and in the cloud. https://aws.amazon.com/security/provable-security/ Key job responsibilities Work with customer teams to understand the nature of their software and the properties they need to establish of it. Identify tools and methods capable of addressing the verification needs of customers, including any novel analysis capabilities required. Use techniques spanning property-based testing to model checkers, and interactive theorem provers to establish program properties. Explore generative AI techniques to help customers formalize their requirements, find revealing tests, generate required boiler plate for testing and model checking, and find and repair program proofs. About the team The Agentic Automated Reasoning Group at AWS develops and applies state of the art formal methods and automated reasoning techniques to ensure the security, reliability, and correctness of AWS services and customer applications, with a strong focus on AI based agents. Our work innovates tools and services to perform verification at scale and apply them to build safe and secure systems at AWS. We are also pioneering the use of formal verification and automated reasoning to develop agentic systems, ensuring AI agents operate within defined safety boundaries.
US, CA, San Francisco
Join the next revolution in robotics at Amazon's Frontier AI & Robotics team, where you'll work alongside world-renowned AI pioneers to lead key initiatives in robotic intelligence. As a Member of Technical Staff, you'll spearhead the development of breakthrough foundation models that enable robots to perceive, understand, and interact with the world in unprecedented ways. You'll drive technical excellence in areas such as perception, manipulation, science understanding, sim2real transfer, multi-modal foundation models, and multi-task learning, designing novel algorithms that bridge the gap between state-of-the-art research and real-world deployment at Amazon scale. In this role, you'll combine hands-on technical work with scientific leadership, ensuring your team delivers robust solutions for dynamic real-world environments. You'll leverage Amazon's vast computational resources to tackle ambitious problems in areas like very large multi-modal robotic foundation models and efficient, promptable model architectures that can scale across diverse robotic applications. Key job responsibilities - Lead technical initiatives in robotics foundation models, driving breakthrough approaches through hands-on research and development in areas like open-vocabulary panoptic scene understanding, scaling up multi-modal LLMs, sim2real/real2sim techniques, end-to-end vision-language-action models, efficient model inference, video tokenization - Design and implement novel deep learning architectures that push the boundaries of what robots can understand and accomplish - Guide technical direction for specific research initiatives, ensuring robust performance in production environments - Mentor and support fellow scientists while maintaining strong individual technical contributions - Collaborate with engineering teams to optimize and scale models for real-world applications - Influence technical decisions and implementation strategies within your area of focus A day in the life - Develop and implement novel foundation model architectures, working hands-on with our extensive compute infrastructure - Guide and support fellow scientists in solving complex technical challenges, from sim2real transfer to efficient multi-task learning - Lead focused technical initiatives from conception through deployment, ensuring successful integration with production systems - Drive technical discussions within your team and with key stakeholders - Conduct experiments and prototype new ideas using our massive compute cluster - Mentor team members while maintaining significant hands-on contribution to technical solutions Amazon offers a full range of benefits that support you and eligible family members, including domestic partners and their children. Benefits can vary by location, the number of regularly scheduled hours you work, length of employment, and job status such as seasonal or temporary employment. The benefits that generally apply to regular, full-time employees include: 1. Medical, Dental, and Vision Coverage 2. Maternity and Parental Leave Options 3. Paid Time Off (PTO) 4. 401(k) Plan If you are not sure that every qualification on the list above describes you exactly, we'd still love to hear from you! At Amazon, we value people with unique backgrounds, experiences, and skillsets. If you’re passionate about this role and want to make an impact on a global scale, please apply! About the team At Frontier AI & Robotics, we're not just advancing robotics – we're reimagining it from the ground up. Our team is building the future of intelligent robotics through ground breaking foundation models and end-to-end learned systems. We tackle some of the most challenging problems in AI and robotics, from developing sophisticated perception systems to creating adaptive manipulation strategies that work in complex, real-world scenarios. What sets us apart is our unique combination of ambitious research vision and practical impact. We leverage Amazon's massive computational infrastructure and rich real-world datasets to train and deploy state-of-the-art foundation models. Our work spans the full spectrum of robotics intelligence – from multimodal perception using images, videos, and sensor data, to sophisticated manipulation strategies that can handle diverse real-world scenarios. We're building systems that don't just work in the lab, but scale to meet the demands of Amazon's global operations. Join us if you're excited about pushing the boundaries of what's possible in robotics, working with world-class researchers, and seeing your innovations deployed at unprecedented scale.
US, WA, Seattle
Innovators wanted! Are you an entrepreneur? A builder? A dreamer? This role is part of an Amazon Special Projects team that takes the company’s Think Big leadership principle to the limits. If you’re interested in innovating at scale to address big challenges in the world, this is the team for you. As an Applied Scientist on our team, you will focus on building state-of-the-art ML models for biology. Our team rewards curiosity while maintaining a laser-focus in bringing products to market. Competitive candidates are responsive, flexible, and able to succeed within an open, collaborative, entrepreneurial, startup-like environment. At the forefront of both academic and applied research in this product area, you have the opportunity to work together with a diverse and talented team of scientists, engineers, and product managers and collaborate with other teams. Key job responsibilities - Build, adapt and evaluate ML models for life sciences applications - Collaborate with a cross-functional team of ML scientists, biologists, software engineers and product managers
US, NY, New York
In this role, you will design and build intelligent multi-agent systems that automate root cause analysis for advertising campaign delivery at scale. You will architect agentic orchestration patterns where specialized sub-agents (campaign diagnostics, deal-level troubleshooting, pacing control) are invoked as composable tools by a reasoning layer that determines which subsystems to query based on the nature of the issue. You will develop hierarchical analysis frameworks that move from daily trend detection to intra-day anomaly isolation, enabling the system to pinpoint when and why delivery degraded rather than relying on static time windows. You will build self-learning feedback loops where the system identifies recurring failure signatures (auction dynamics, pacing anomalies, supply contention), updates its diagnostic knowledge as engineering teams deploy fixes, and retires stale patterns automatically. We are looking for a passionate Applied Scientist with technical expertise in LLM-based agent architectures, retrieval-augmented generation, time-series anomaly detection, and production ML systems. In addition to hands-on experience building agentic AI solutions, an ideal candidate should demonstrate the ability to translate complex distributed system behaviors into structured diagnostic reasoning, show a willingness to push the boundaries of how LLMs interact with real-time operational data, and thrive in an environment where you ship production systems that directly reduce advertiser escalation time from days to minutes. Key job responsibilities * Conduct deep data analysis to derive insights for the business, identify gaps, and uncover new opportunities. * Develop scalable and effective machine learning models and optimization strategies to solve business problems. * Run regular A/B experiments, gather data, and perform statistical analysis to optimize advertiser experiences. * Collaborate closely with software engineers to deliver end-to-end solutions into production. * Enhance the scalability, efficiency, and automation of large-scale data analytics, model training, deployment, and serving. * Research and implement new machine learning models and techniques to improve advertising performance. A day in the life Your primary focus is building a multi-agent diagnostic system that automates root cause analysis for advertising campaign delivery issues. On a typical day, you might review how the system handled recent escalations, identify where it reasoned incorrectly, adjust orchestration logic, and write new evaluation cases. You will design agent architectures that invoke specialized sub-agents as tools, build hierarchical analysis frameworks that move from trend detection to anomaly isolation, and develop self-learning loops that keep the system's diagnostic knowledge current as the underlying platform evolves. You will work closely with SDEs building the diagnostic platform, product managers defining the troubleshooting experience, and the support teams who rely on your system to resolve advertiser delivery issues in minutes instead of days. Beyond the core agent work, you may find yourself diving into causal inference to measure recommendation effectiveness, prototyping proactive anomaly detection, or contributing to evaluation science for systems that reason over complex operational data. About the team The Demand Enablement, Product Analytics and Operations team builds the diagnostic and intelligence layer for Amazon DSP, the demand-side platform powering Amazon's programmatic advertising business. We own the systems that detect, diagnose, and surface delivery issues across campaigns, giving internal teams and advertisers the visibility to act before problems impact spend. Our product portfolio spans automated troubleshooting platforms, advertiser-facing delivery insights, and AI-powered root cause analysis using multi-agent architectures on foundation models. We are a small, high-ownership team that ships production systems end-to-end, from data pipelines processing billions of bid events to LLM-based agents that reason over complex advertising systems. If you want to work at the intersection of applied science, distributed systems observability, and real business impact measured in advertiser dollars recovered, this is the team.