How We Make Alexa Sound More Human-like Lead.png
Andrew Breen, senior manager, Amazon text-to-speech research, provided an overview of the history of TTS advancements at last June's re:MARS conference.

Advances in text-to-speech technologies help computers find their voice

Generating natural sounding, human-like speech has been a goal of scientists for decades.

Editor's Note: The Alexa team recently introduced a new longform speaking style so Alexa sounds more natural when reading long pieces of content, like this article. If you prefer to listen to this story rather than read it, below is this article utilizing the longform speaking style.

The spoken word is important to people. We love the sound of our child’s voice, of a favorite song, or of our favorite movie star reciting a classic line.

Computer-generated, synthesized spoken words also are becoming increasingly common. Alexa, Amazon’s popular voice service, has been responding to customers’ questions and requests for more than five years, and is now available on hundreds of millions of devices from Amazon and third-party device manufacturers. Other businesses also are taking advantage of computer-generated speech to handle customer service calls, market products, and more.

How we make Alexa sound more human-like

Language and speech are incredibly complex. Words have meaning, sure. So does the context of those words, the emotion behind them, and the response of the person listening. It would seem the subtleties of the spoken word would be beyond the reach of even the most sophisticated computers. But in recent years, advances in text-to-speech (TTS) technologies – the ability of computers to convert sequences of words into natural sounding, intelligible audio responses – have made it possible for computers to sound more human-like.

Amazon scientists and engineers are helping break new ground in an era where computers sound not only friendly and knowledgeable, but also predict how the sentiment of an utterance might sound to an average listener, for example, and respond with human-like intonations.

A revolution within the field occurred in 2016, when WaveNet – a technology for generating raw audio – was introduced. Created by researchers at London-based artificial intelligence firm DeepMind, the technique could generate realistic voices using a neural network trained with recordings of real speech.

Andrew Breen (crop)
Andrew Breen, senior manager, TTS research

“This early research suggested that a new machine learning method offered equal or greater quality and the potential for more flexibility,” says Andrew Breen, senior manager of the TTS research team in Cambridge, UK. Breen has long worked on the problem of making computerized speech more responsive and authentic. Before joining Amazon in 2018, he was director of TTS research for Nuance, a Massachusetts-based company that develops conversational artificial intelligence solutions.

Modeled loosely on the human neural system, neural nets are networks of simple but densely interconnected processing nodes. Typically, those nodes are arranged into layers, and the output of each layer passes to the layer above it. The connections between layers have associated “weights” that determine how much the output of one node contributes to the computation performed by the next.

Combined with machine learning, neural networks have accelerated progress in improving computerized speech. “It’s really a gold rush of invention,” says Breen.

Generating natural-sounding speech

Generating natural sounding, human-like speech has been a goal of scientists for decades. In the 1930s Bell Labs scientist Homer Dudley developed the Voder, a primitive synthetic-speech machine that an operator worked like a piano keyboard – except rather than music, out came a squawking mechanical voice. In the 1980s, a computerized TTS application called DECTalk, developed by the Digital Equipment Corporation, had progressed to the point where the late Stephen Hawking could use a version of it, paired with a keyboard to “talk”. The results were artificial-sounding, but intelligible words that many people still associated with a talking machine.

It's really been a gold rush of invention.
Andrew Breen, senior manager, TTS research

By the early 2000s, more accurate speech synthesis became common. The foremost approach taken then: hybrid unit concatenation. Amazon, for instance, used this approach until 2015 to build early versions of Alexa’s voice or to build voice capabilities into products like the Fire Tablet. Says Nikhil Sharma, a principal product manager in Amazon’s TTS group: “To create some of the early Alexa voices, we worked with voice talents in a studio for hours and had them say a wide variety of phrases. We broke that speech data down into a single diphone (a single diphone is a combination of halves of two phonemes, a distinct unit of sound) and put that in a large audio database. Then, when a request came to generate speech, we could tap into that database and select the best diphones to stitch together and create a sentence spoken by Alexa.”

Nikhil Sharma, principal product manager, TTS, Amazon
Nikhil Sharma, principal product manager, TTS

That process worked fairly well. But hybrid unit concatenation has its limits. It needs large amounts of pre-recorded sounds from professional voice talent for reference – sort of like a tourist constantly flipping through a large French book to find particular phrases. “Because of that, we really couldn’t say a hybrid unit concatenation system ‘learned’ a language,” says Breen.

Creating a computer that actually learns a language – not just memorizes phrases – became a goal of researchers. “That has been the Holy Grail, but nobody knew how to do it,” says Breen. “We were close but had a quality ceiling that limited its viability.”

Neural networks offered a way to do just that. In 2018, Amazon scientists demonstrated that by using a generative neural network approach to creating synthetic speech, they could produce natural sounding speech. Using the generative neural network approach, Alexa could also flex the way she speaks about certain content. For example, Amazon scientists created Alexa’s newscaster style of speech from just a few hours of training data, allowing customers to hear the news in a style to which they’ve become accustomed. This advance paved the way for Alexa and other Amazon services to adopt different speaking styles in different contexts, improving customer experiences.

Comparisons of Alexa synthesized speech

Star Trek

2014 Concatenative
2020 NTTS

Song ID

Standard response
Music style response
Above are examples of how Alexa's voice has become more natural over the years.

Amazon recently announced a new Amazon Polly feature called Brand Voice, which provides the opportunity for organizations to work with the Amazon Polly team of AI research scientists and linguists to build an exclusive, high-quality, neural TTS voice that represents their brand’s persona. Early adopters Kentucky Fried Chicken (KFC) Canada and National Australia Bank (NAB) have utilized the service to each create two unique brand voices that utilize the same deep learning technology that powers the voice of Alexa.

Amazon Polly is an AWS service that turns text into lifelike speech, allowing customers to build entirely new categories of speech-enabled products. Polly provides dozens of lifelike voices across a broad set of languages, allowing customers to build speech-enabled applications that work in many different countries.

Looking forward, Amazon researchers are working toward teaching computers to understand the meaning of a set of words, and speak those words using the appropriate affect. “If I gave a computer a news article, it would do a reasonable job of rendering the words in the article,” says Breen. “But it’s missing something. What is missing is the understanding of what is in the article, whether it’s good news or bad, and what is the focal point. It lacks that intuition.”

That is changing. Now, computers can be taught to say the same sentence with varying kinds of inflection. In the future, it’s possible they’ll recognize how they should be saying those words based simply on the context of the words, or the words themselves. “We want computers to be sensitive to the environment and to the listener, and adapt accordingly,” says Breen.

There are numerous potential TTS applications, from customer service and remote learning to narration of news articles. Driving improvements in this technology is one approach Amazon scientists and engineers are taking to create better experiences, not only for Alexa customers, but for organizations worldwide.

“The ability for Alexa to adapt her speaking style based on the context of a customer's request opens the possibility to deliver new and delightful experiences that were previously unthinkable,” says Breen. “These are really exciting times.”

Related content

US, CA, Palo Alto
The Amazon Search team creates powerful, customer-focused search and advertising solutions and technologies. Whenever a customer visits an Amazon site worldwide and types in a query or browses through product categories, the Amazon Search services go to work. We design, develop, and deploy high performance, fault-tolerant distributed search systems used by millions of Amazon customers every day. Our team works to maximize the quality and effectiveness of the search experience for visitors to Amazon websites worldwide.
JP, Tokyo
The Amazon Logistics (AMZL) Team is responsible for the acquisition, design, construction, and management of all facilities in the Amazon Delivery Station Network. AMZL is looking for a talented and passionate Data Scientist to help shape its Last Mile business with technical strategies and solutions, by processing, analyzing and interpreting huge data sets. You should be comfortable with ambiguity, problem solving and enjoy working in a fast-paced, diverse and dynamic environment. Using analytical rigor and statistical methods, you mine through data to identify opportunities for Amazon and our delivery channels. And you collaborate with other scientists, engineers, Product and Program Managers to deploy new products and solutions. [More Information] Last Mile Department Data Analyst/BI Engineer Tokyo Office *Amazon is committed to a diverse and inclusive workplace. Amazon is an equal opportunity employer and does not discriminate on the basis of race, national origin, gender, gender identity, sexual orientation, protected veteran status, disability, age, or other legally protected status. For individuals with disabilities who would like to request an accommodation, visit https://www.amazon.jobs/disability/jp Key job responsibilities Creating a roadmap of the most challenging business questions and use data to articulate possible root cause analysis and solutions Managing and executing entire projects or components of large projects from start to finish including project management, data gathering and manipulation, synthesis and modeling, problem solving, and communication of insights Partnering with Product, Program and Engineering teams to design and run models, research new algorithms, and prove incrementality and drive growth Understanding drivers, impacts, and key influences on seller growth dynamics Developing and scaling end-to-end ML Models and solutions Automating feedback loops for algorithms in production Utilizing Amazon systems and tools to effectively work with terabytes of data About the team Last Mile Execution Analytics (LMEA) team of JP works as an integral part of Amazon Logistics to ensure that its business intelligence, analytics, tools and planning needs are met. By providing information, insight, and decision support, we strive to enable success of all parts of AMZL. Our customer set includes senior management, station operations, external vendors, long-term planning, Ops technology (Voice of the Delivery Station, Voice of the Customer), network planning, and pretty much every BI and Ops teams. Voice of Employee [Work Life Harmony] We believe, it is important to spend private time such as spending time with your family or doing anything you like to spur innovation. Amazon promotes a fulfilling and flexible work style according to the work volume and lifestyle of each employee.
US, CA, San Francisco
About Twitch Launched in 2011, Twitch is a global community that comes together each day to create multiplayer entertainment: unique, live, unpredictable experiences created by the interactions of millions. We bring the joy of co-op to everything, from casual gaming to world-class esports to anime marathons, music, and art streams. Twitch also hosts TwitchCon, where we bring everyone together to celebrate, learn, and grow their personal interests and passions. We’re always live at Twitch. Stay up to date on all things Twitch on Linkedin, Twitter and on our Blog. About the role: Twitch builds data-driven machine learning solutions across several rich problem spaces: Natural Language Processing (NLP), Recommendations, Semantic Search, Classification/Categorization, Anomaly Detection, Forecasting, Safety, and HCI/Social Computing/Computational Social Science. As an Intern, you will work with a dedicated Mentor and Manager on a project in one of these problem areas. You will also be supported by an Advisor and participate in cohort activities such as research teach backs and leadership talks. This position can also be located in San Francisco, CA or virtual. You Will: Solve large-scale data problems. Design solutions for Twitch's problem spaces Explore ML and data research
US, CA, San Francisco
About Twitch Launched in 2011, Twitch is a global community that comes together each day to create multiplayer entertainment: unique, live, unpredictable experiences created by the interactions of millions. We bring the joy of co-op to everything, from casual gaming to world-class esports to anime marathons, music, and art streams. Twitch also hosts TwitchCon, where we bring everyone together to celebrate, learn, and grow their personal interests and passions. We’re always live at Twitch. Stay up to date on all things Twitch on Linkedin, Twitter and on our Blog. About the role: Twitch builds data-driven machine learning solutions across several rich problem spaces: Natural Language Processing (NLP), Recommendations, Semantic Search, Classification/Categorization, Anomaly Detection, Forecasting, Safety, and HCI/Social Computing/Computational Social Science. As an Intern, you will work with a dedicated Mentor and Manager on a project in one of these problem areas. You will also be supported by an Advisor and participate in cohort activities such as research teach backs and leadership talks. This position can also be located in San Francisco, CA or virtual. You Will: Solve large-scale data problems. Design solutions for Twitch's problem spaces Explore ML and data research
LU, Luxembourg
Are you a talented and inventive scientist with a strong passion about modern data technologies and interested to improve business processes, extracting value from the data? Would you like to be a part of an organization that is aiming to use self-learning technology to process data in order to support the management of the procurement function? The Global Procurement Technology, as a part of Global Procurement Operations, is seeking a skilled Data Scientist to help build its future data intelligence in business ecosystem, working with large distributed systems of data and providing Machine Learning (ML) and Predictive Modeling expertise. You will be a member of the Data Engineering and ML Team, joining a fast-growing global organization, with a great vision to transform the Procurement field, and become the role model in the market. This team plays a strategic role supporting the core Procurement business domains as well as it is the cornerstone of any transformation and innovation initiative. Our mission is to provide a high-quality data environment to facilitate process optimization and business digitalization, on a global scale. We are supporting business initiatives, including but not limited to, strategic supplier sourcing (e.g. contracting, negotiation, spend analysis, market research, etc.), order management, supplier performance, etc. We are seeking an individual who can thrive in a fast-paced work environment, be collaborative and share knowledge and experience with his colleagues. You are expected to deliver results, but at the same time have fun with your teammates and enjoy working in the company. In Amazon, you will find all the resources required to learn new skills, grow your career, and become a better professional. You will connect with world leaders in your field and you will be tackling Data Science challenges to ensure business continuity, by taking the right decisions for your customers. As a Data Scientist in the team, you will: -be the subject matter expert to support team strategies that will take Global Procurement Operations towards world-class predictive maintenance practices and processes, driving more effective procurement functions, e.g. supplier segmentation, negotiations, shipping supplies volume forecast, spend management, etc. -have strong analytical skills and excel in the design, creation, management, and enterprise use of large data sets, combining raw data from different sources -provide technical expertise to support the development of ML models to facilitate intelligent digital services, such as Contract Lifecycle Management (CLM) and Negotiations platform -cooperate closely with different groups of stakeholders, e.g. data/software engineers, product/program managers, analysts, senior leadership, etc. to evaluate business needs and objectives to set up the best data management environment -create and share with audiences of varying levels technical papers and presentations -deal with ambiguity, prioritizing needs, and delivering results in a dynamic environment Basic qualifications -Master’s Degree in Computer Science/Engineering, Informatics, Mathematics, or a related technical discipline -3+ years of industry experience in data engineering/science, business intelligence or related field -3+ years experience in algorithm design, engineering and implementation for very-large scale applications to solve real problems -Very good knowledge of data modeling and evaluation -Very good understanding of regression modeling, forecasting techniques, time series analysis, machine-learning concepts such as supervised and unsupervised learning, classification, random forest, etc. -SQL and query performance tuning skills Preferred qualifications -2+ years of proficiency in using R, Python, Scala, Java or any modern language for data processing and statistical analysis -Experience with various RDBMS, such as PostgreSQL, MS SQL Server, MySQL, etc. -Experience architecting Big Data and ML solutions with AWS products (Redshift, DynamoDB, Lambda, S3, EMR, SageMaker, Lex, Kendra, Forecast etc.) -Experience articulating business questions and using quantitative techniques to arrive at a solution using available data -Experience with agile/scrum methodologies and its benefits of managing projects efficiently and delivering results iteratively -Excellent written and verbal communication skills including data visualization, especially in regards to quantitative topics discussed with non-technical colleagues
US, CA, San Francisco
About Twitch Launched in 2011, Twitch is a global community that comes together each day to create multiplayer entertainment: unique, live, unpredictable experiences created by the interactions of millions. We bring the joy of co-op to everything, from casual gaming to world-class esports to anime marathons, music, and art streams. Twitch also hosts TwitchCon, where we bring everyone together to celebrate, learn, and grow their personal interests and passions. We’re always live at Twitch. Stay up to date on all things Twitch on Linkedin, Twitter and on our Blog. About the role: Twitch builds data-driven machine learning solutions across several rich problem spaces: Natural Language Processing (NLP), Recommendations, Semantic Search, Classification/Categorization, Anomaly Detection, Forecasting, Safety, and HCI/Social Computing/Computational Social Science. As an Intern, you will work with a dedicated Mentor and Manager on a project in one of these problem areas. You will also be supported by an Advisor and participate in cohort activities such as research teach backs and leadership talks. This position can also be located in San Francisco, CA or virtual. You Will: Solve large-scale data problems. Design solutions for Twitch's problem spaces Explore ML and data research
US, CA, San Francisco
About Twitch Launched in 2011, Twitch is a global community that comes together each day to create multiplayer entertainment: unique, live, unpredictable experiences created by the interactions of millions. We bring the joy of co-op to everything, from casual gaming to world-class esports to anime marathons, music, and art streams. Twitch also hosts TwitchCon, where we bring everyone together to celebrate, learn, and grow their personal interests and passions. We’re always live at Twitch. Stay up to date on all things Twitch on Linkedin, Twitter and on our Blog. About the role: Twitch builds data-driven machine learning solutions across several rich problem spaces: Natural Language Processing (NLP), Recommendations, Semantic Search, Classification/Categorization, Anomaly Detection, Forecasting, Safety, and HCI/Social Computing/Computational Social Science. As an Intern, you will work with a dedicated Mentor and Manager on a project in one of these problem areas. You will also be supported by an Advisor and participate in cohort activities such as research teach backs and leadership talks. This position can also be located in San Francisco, CA or virtual. You Will: Solve large-scale data problems. Design solutions for Twitch's problem spaces Explore ML and data research
US, CA, San Francisco
About Twitch Launched in 2011, Twitch is a global community that comes together each day to create multiplayer entertainment: unique, live, unpredictable experiences created by the interactions of millions. We bring the joy of co-op to everything, from casual gaming to world-class esports to anime marathons, music, and art streams. Twitch also hosts TwitchCon, where we bring everyone together to celebrate, learn, and grow their personal interests and passions. We’re always live at Twitch. Stay up to date on all things Twitch on Linkedin, Twitter and on our Blog. About the role: Twitch builds data-driven machine learning solutions across several rich problem spaces: Natural Language Processing (NLP), Recommendations, Semantic Search, Classification/Categorization, Anomaly Detection, Forecasting, Safety, and HCI/Social Computing/Computational Social Science. As an Intern, you will work with a dedicated Mentor and Manager on a project in one of these problem areas. You will also be supported by an Advisor and participate in cohort activities such as research teach backs and leadership talks. This position can also be located in San Francisco, CA or virtual. You Will: Solve large-scale data problems. Design solutions for Twitch's problem spaces Explore ML and data research
US, CA, San Francisco
About Twitch Launched in 2011, Twitch is a global community that comes together each day to create multiplayer entertainment: unique, live, unpredictable experiences created by the interactions of millions. We bring the joy of co-op to everything, from casual gaming to world-class esports to anime marathons, music, and art streams. Twitch also hosts TwitchCon, where we bring everyone together to celebrate, learn, and grow their personal interests and passions. We’re always live at Twitch. Stay up to date on all things Twitch on Linkedin, Twitter and on our Blog. About the role: Twitch builds data-driven machine learning solutions across several rich problem spaces: Natural Language Processing (NLP), Recommendations, Semantic Search, Classification/Categorization, Anomaly Detection, Forecasting, Safety, and HCI/Social Computing/Computational Social Science. As an Intern, you will work with a dedicated Mentor and Manager on a project in one of these problem areas. You will also be supported by an Advisor and participate in cohort activities such as research teach backs and leadership talks. This position can also be located in San Francisco, CA or virtual. You Will: Solve large-scale data problems. Design solutions for Twitch's problem spaces Explore ML and data research
US, WA, Seattle
We are a team of doers working passionately to apply cutting-edge advances in deep learning in the life sciences to solve real-world problems. As a Senior Applied Science Manager you will participate in developing exciting products for customers. Our team rewards curiosity while maintaining a laser-focus in bringing products to market. Competitive candidates are responsive, flexible, and able to succeed within an open, collaborative, entrepreneurial, startup-like environment. At the leading edge of both academic and applied research in this product area, you have the opportunity to work together with a diverse and talented team of scientists, engineers, and product managers and collaborate with others teams. Location is in Seattle, US Embrace Diversity Here at Amazon, we embrace our differences. We are committed to furthering our culture of inclusion. We have ten employee-led affinity groups, reaching 40,000 employees in over 190 chapters globally. We have innovative benefit offerings, and host annual and ongoing learning experiences, including our Conversations on Race and Ethnicity (CORE) and AmazeCon (gender diversity) conferences. Amazon’s culture of inclusion is reinforced within our 14 Leadership Principles, which remind team members to seek diverse perspectives, learn and be curious, and earn trust Balance Work and Life Our team puts a high value on work-life balance. It isn’t about how many hours you spend at home or at work; it’s about the flow you establish that brings energy to both parts of your life. We believe striking the right balance between your personal and professional life is critical to life-long happiness and fulfillment. We offer flexibility in working hours and encourage you to find your own balance between your work and personal lives Mentor & Grow Careers Our team is dedicated to supporting new members. We have a broad mix of experience levels and tenures, and we’re building an environment that celebrates knowledge sharing and mentorship. Our senior members enjoy one-on-one mentoring and thorough, but kind, code reviews. We care about your career growth and strive to assign projects based on what will help each team member develop into a better-rounded engineer and enable them to take on more complex tasks in the future. Key job responsibilities • Manage high performing engineering and science teams • Hire and develop top-performing engineers, scientists, and other managers • Develop and execute on project plans and delivery commitments • Work with business, data science, software engineer, biological, and product leaders to help define product requirements and with managers, scientists, and engineers to execute on them • Build and maintain world-class customer experience and operational excellence for your deliverables