Prem Natarajan, Alexa AI vice president of natural understanding, giving a presentation
Prem Natarajan, Alexa AI vice president of natural understanding
Credit: Micron Technology, Inc.

3 questions: Prem Natarajan on issues of AI fairness and bias

Alexa AI vice president of natural understanding Prem Natarajan discusses the upcoming cycle for the National Science Foundation collaboration on fairness in AI, his participation on the Partnership on AI board, and issues related to bias in natural language processing.

A year ago, Amazon and the National Science Foundation (NSF) announced a $20 million collaboration to fund academic research on fairness in AI over a three-year period. Recently, Erwin Gianchandani, deputy assistant director for Computer and Information Science and Engineering at NSF, discussed the work of the first ten recipients of the program’s grants. Here, Prem Natarajan, Alexa AI vice president of natural understanding, and the Amazon executive who helped launch the collaboration with NSF, discusses the next cycle of upcoming proposals from academic researchers, his work with the Partnership on AI, and what can be done to address bias in natural language processing models.

The 2020 award cycle for the Fairness in AI program in conjunction with the NSF recently launched. Full proposals are due by July 13th. What are you hoping to see in the next round of proposals?

We collaborated with the NSF to launch the Fairness in AI program with the goal of promoting academic research in this important aspect of AI. Our primary objective for engaging with academia on issues related to fairness and transparency in AI is to get many different and diverse perspectives focused on the challenge. The teams selected by NSF in the first round are addressing a variety of topics – from principled frameworks for developing and certifying fair AI, to domain-focused applications such as fair recommender systems for foster care services. To that end, I hope that the second round will build upon the success of the first round by bringing an even greater diversity of perspectives on definitions and perceptions of fairness. Without such diversity the entire field of research into fair AI will become a self-defeating exercise.

Another hope I have for the second round, and indeed for all rounds of this program, is that it will drive the creation of a portfolio of open-source artifacts – such as data sets, metrics, tools, and testing methodologies – which all stakeholders in AI can use to promote the use of fair AI. Such readily available artifacts will make it easier for the community to learn from one another, promote the replication of research results, and, ultimately, advance the state of the art more rapidly. Put differently, we hope that open access to the research under this program will form a rising tide that lifts all boats. It also seems natural that methodologies for fairness will benefit from broad and inclusive discussion across relevant academic and scientific communities.

The deadline for this next round of proposal submissions is July 13th. We hope that the response to this round will be even stronger than for the first. NSF selects the recipients, and I am sure NSF’s reviewers are looking forward to a summer of interesting reading!

You are Amazon’s representative on the Partnership on AI (PAI) board of directors. This unique organization has thematic pillars related to safety-critical AI; fair, transparent and accountable AI; AI labor and the economy; collaborations between AI systems and people; social and societal influences of AI; and AI and social good. It’s an ambitious, broad agenda. You’re fairly new in your role with PAI; what most excites you about the work being done there?

The most exciting aspect of the Partnership on AI is that it is a unique multi-sector forum where I get to listen to and learn from the incredible diversity of perspectives – from industry, academia, non-profits, and social justice groups. PAI today counts amongst its members about 59 non-profits, 24 academic institutions, and 18 industrial organizations. While I joined the board just a few months ago, I have already attended several meetings and participated in discussions with other PAI members as well as PAI staff. While every member has their own unique perspective on AI, it’s been really interesting and encouraging to see that we all share the same values and many of the same concerns. It should be of no surprise that the issue of equity is top of mind with a concomitant focus on fairness considerations.

Alexa & Friends Twitch show features Prem Natarajan

Earlier this month, Alexa evangelist Jeff Blankenburg interviewed Prem Natarajan live on the 'Alexa & Friends' Twitch show. In the video, they discuss recent advances in natural understanding , and how those advancements translate into better experiences for customers, developers and third-party device manufacturers.

From a technical perspective, I am excited by the number and quality of research initiatives underway at PAI. Many of these initiatives are of critical importance to the future development of the field of AI. Let me give you a couple of examples.

One is the area of fairness, accountability and transparency. There are several projects underway in this area, but I will mention one that to me exemplifies the kind of work that an organization like PAI can do. PAI researchers interviewed practitioners at twenty different organizations and performed an in-depth case study of how explainable AI is used today. This kind of research is very important to AI practitioners because it gives them a referential basis to assess their own work and to identify useful areas for future contributions.

Another example is ABOUT ML, which is focused on developing and sharing best practices as well as on advancing public understanding of AI. A couple of years ago some researchers had proposed the development of an AI model scorecard, along the lines of the nutritional information you get on the back of most food items we buy today. The scorecard would describe the attributes of the data used to train the models, the way in which it was tested, etc. The motivation behind the scorecard is to give other developers or model builders a sense of the strengths and limitations of the model, so they can better estimate and address potential weaknesses in the model for their target use cases. ABOUT ML goes well beyond such a scorecard, focusing on documentation, provenance of data and code artifacts, and other critical attributes of the model development process. Ultimately, only multisector organizations like PAI can successfully drive this kind of initiative, bringing together people across organizations and sectors.

Lastly, there’s an education role that PAI serves that I believe is unique, serving as the bridge between AI technologists and other stakeholders within society, making sure AI technologists are appropriately factoring in the perspectives and concerns of the other stakeholders within society. Some examples here include PAI’s collaborative work with First Draft, a PAI Partner, to help technologists and journalists at digital platforms address growing issues around manipulated media. PAI also helps those stakeholders understand more about how AI technology works, its strengths and its limitations.

You oversee Alexa’s natural understanding team. Natural language processing models have drawn criticism for capturing common social biases with respect to gender and race. A large body of work is emerging related to bias in word embedding and classifiers, and there are many proposals for countermeasures. Can you describe the challenge of bias in NLP models, and give us insight into some of the countermeasures you think are, or could be, effective?

A word embedding is a vector of real numbers representing that word; the core idea is that words with similar meanings map to vectors that are “close” to each other. Word embeddings have become a central feature of modern NLP. While embeddings can be computed using a variety of different techniques, deep learning techniques have proven to be tremendously effective at numerically representing the semantics of a word and concepts, etc. Today, deep learning based embeddings are used for all kinds of processing, from named entity recognition, to question answering, and natural language generation. As a result, the semantics that these embeddings encode greatly influence how we interpret text, the accuracy of those interpretations, and the actions we take in response to those interpretations.

Bias can also manifest in other ways because any system that is based on data can exhibit a majoritarian bias to it.
Prem Natarajan, Alexa AI VP of natural understanding

As word embeddings became prevalent, researchers naturally started looking into their fragilities and shortcomings. One of those fragilities is that the embeddings derive and encode meaning from context, which means that the meaning of a word is largely controlled by the different contexts in which that word is observed in the training data. While that seems like a reasonable basis for inferring meaning, it leads to undesirable consequences. My friend Kai-Wei Chang at UCLA is one of the early investigators of bias in NLP and he uses the following example: take the vector for doctor and you subtract the vector for man; when you add the vector for woman, you should in principle get the vector for doctor again, or a female doctor. But instead the resulting vector is close to the vector for ‘nurse.’ What this example shows is that the latent biases in human-generated text get encoded into the embeddings. One example of a system that is affected by these biases is natural language generation. Many studies have shown that such biases can result in the generation of text that exhibits the same biases and prejudices as humans, sometimes in an amplified manner. Left unmitigated, such systems could reinforce human biases and stereotypes.

Bias can also manifest in other ways because any system that is based on data can exhibit a majoritarian bias to it. So, for example, different groups in different parts of the world may speak the same language with different dialects, but the most frequent dialect will likely see the best performance only because it forms the major proportion of the training data. But we don’t want dialect or accent to determine how well the system will work for an individual. We want our systems to work equally well for everyone, regardless of geography, dialect, gender, or any other irrelevant factor.

Methodologically, we counter the impact of bias by using a principled approach to characterize the dimensions of bias and associated impact, and by developing techniques that are robust to these biasing factors. For example, it stands to reason that speech recognition systems should ignore parts of the signal that are not useful for recognizing the words that were spoken. It shouldn’t really matter whether the voice is male or female, only the actual words should. Similarly for natural language understanding, we want to be able to understand the queries of different groups of people regardless of the stylistic or syntactic variations of the language used. Scientists at Amazon and elsewhere are exploring a broad variety of approaches such as de-biasing techniques, adversarial invariance, active learning, and selective sampling. Personally, I find the adversarial approaches to both testing and to generating bias or nuisance invariant representations most appealing because of their scalability, but in the next few years, we will all find out what works best for different problems!

Research areas

Related content

  • Staff writer
    April 14, 2026
    Built in collaboration with the Gray Lab at Johns Hopkins Whiting School of Engineering, the Antibody Developability Benchmark is powered by one of the most diverse antibody datasets in public literature, enabling transparent performance evaluation for AI-guided antibody design.
  • Meiqi Sun
    April 20, 2026
    Large language models today can solve algebra, pass academic benchmarks, and generate highly structured chain-of-thought explanations. In text-only settings, they often feel startlingly intelligent — methodical, articulate, even strategic. But place those models inside an interactive environment — ask them to click buttons, scroll pages, fill out forms, and submit answers — and their behavior changes. Their careful reasoning falters. They guess where they once deduced. They adhere to templates and produce limited procedural narration: stating what they see and what they will click next, without first forming a structured plan and acting in accordance with plan. It’s as if part of their intelligence has quietly gone offline the moment the cursor appears.
    Machine learning
  • Louise Ping, John Gray, Emily Webber, Josh Longenecker
    August 10, 2026
    A competition with a finalist ceremony during NeurIPS 2026, challenging researchers to train language models from scratch on Trainium, exploring what optimal architectures look like when the hardware changes.
US, WA, Seattle
This role sits within Amazon's Automated Reasoning and Formal Verification research horizon. Shape the Future of Cloud Computing. Are you a graduate student passionate about Automated Reasoning and its real-world applications? Join our team of innovators and embark on a journey to revolutionize cloud computing through innovative automated reasoning techniques. Our tools are called billions of times daily, powering the backbone of Amazon's products and services. We are changing the way computer systems are developed and operated, raising the bar for security, durability, availability, and quality. Applied Scientists in Automated Reasoning develop and apply formal methods, automated reasoning techniques, and neurosymbolic approaches to ensure the security, reliability, and correctness of Amazon and AWS services and customer applications. Application areas span cloud infrastructure verification, cryptographic assurance, AI safety, and formal guarantees for generative AI systems. Methods range from interactive theorem proving and constraint solving to neuro-inspired proof search. As an Applied Science Intern, you will have the opportunity to work alongside our scientists and contribute to projects. From distributed proof search and SAT/SMT solvers to program analysis, synthesis, and verification, you will tackle complex challenges at the intersection of theory and practice. Amazon has positions available for Automated Reasoning Applied Science Internships in, but not limited to, Arlington, VA; Boston, MA; New York, NY; Portland, OR; Santa Clara, CA; Seattle, WA. Key job responsibilities We are particularly interested in candidates with expertise in: Theorem Proving, Boolean Satisfiability Solvers, Bounded Model Checking, Deductive Verification, Programming/Scripting Languages, Abstract Interpretation, Automated Reasoning, Static/Program Analysis, Program Synthesis. Contribute to the design and implementation of algorithms and formal methods for automated reasoning, including constraint solving, model checking, static analysis, theorem proving, and program synthesis, within a guided research framework. Explore and apply generative AI and machine learning techniques to enhance automated reasoning, including learning-based heuristics for search, neural approaches to symbolic reasoning, and methods for verifying the correctness of AI-generated code. Contribute to automated reasoning techniques for generative AI and agentic coding systems, including methods that apply formal guarantees to large language model outputs. Contribute to the scientific community through publications at peer-reviewed conferences and journals. Leverage AI-powered tools where applicable to accelerate research, experimentation, and prototyping. Critically review and validate outputs from AI tools and automated systems. The ideal intern must have the ability to communicate research findings clearly to diverse audiences.
US, NY, New York
We are seeking an Applied Scientist to lead the development of evaluation frameworks and data collection protocols for robotic capabilities. In this role, you will focus on designing how we measure, stress-test, and improve robot behavior across a wide range of real-world tasks. Your work will play a critical role in shaping how policies are validated and how high-quality datasets are generated to accelerate system performance. You will operate at the intersection of robotics, machine learning, and human-in-the-loop systems, building the infrastructure and methodologies that connect teleoperation, evaluation, and learning. This includes developing evaluation policies, defining task structures, and contributing to operator-facing interfaces that enable scalable and reliable data collection. The ideal candidate is highly experimental, systems-oriented, and comfortable working across software, robotics, and data pipelines, with a strong focus on turning ambiguous capability goals into measurable and actionable evaluation systems. Key job responsibilities - Design and implement evaluation frameworks to measure robot capabilities across structured tasks, edge cases, and real-world scenarios - Develop task definitions, success criteria, and benchmarking methodologies that enable consistent and reproducible evaluation of policies - Create and refine data collection protocols that generate high-quality, task-relevant datasets aligned with model development needs - Build and iterate on teleoperation workflows and operator interfaces to support efficient, reliable, and scalable data collection - Analyze evaluation results and collected data to identify performance gaps, failure modes, and opportunities for targeted data collection - Collaborate with engineering teams to integrate evaluation tooling, logging systems, and data pipelines into the broader robotics stack - Stay current with advances in robotics, evaluation methodologies, and human-in-the-loop learning to continuously improve internal approaches - Lead technical projects from conception through production deployment - Mentor junior scientists and engineers About the team Fauna Robotics, an Amazon company, is building capable, safe, and genuinely delightful robots for everyday life. Our goal is simple: make robots people actually want to live and interact with in everyday human spaces. We believe that future won’t arrive until building for robotics becomes far more accessible. Today, too much effort is spent reinventing the fundamentals. We’re changing that by developing tightly integrated hardware and software systems that make it faster, safer, and more intuitive to create real-world robotic products. Our work spans the full stack: mechanical design, control systems, dynamic modeling, and intelligent software. The focus is not just functionality, but experience. We’re building robots that feel responsive, expressive, and genuinely useful. At Fauna, you’ll work at the frontier of this space, helping define how robots move, manipulate, and interact with people in natural environments. It’s an opportunity to solve hard problems across hardware and software with a team focused on making robotics accessible and joyful to build. If you care about making robotics real for everyone and building systems that are as delightful as they are capable, we’re interested in hearing from you.
US, WA, Seattle
We are looking for a talented, organized, and customer-focused applied researcher to join our Pricing Optimization science group, with a charter to measure, refine, and launch customer-obsessed improvements to our algorithmic pricing and promotion models across all products listed on Amazon. This role requires an individual with exceptional machine learning modeling and architecture expertise — particularly in deep learning, neural networks, and transformer-based architectures applied to price prediction and forecasting problems. Equally important is deep expertise in causal machine learning — including causal inference, treatment-effect estimation, and experimentation methods (e.g., uplift modeling, double/debiased machine learning, instrumental variables, and A/B and quasi-experimental design) — to isolate the true impact of pricing and promotion decisions on customer behavior and business outcomes. The ideal candidate brings a strong foundation in applied statistics and probabilistic modeling, excellent cross-functional collaboration skills, business acumen, and an entrepreneurial spirit. We are looking for an experienced innovator who is a self-starter, comfortable with ambiguity, demonstrates strong attention to detail, and has the ability to work in a fast-paced and ever-changing environment. Key job responsibilities See the big picture. Understand and influence the long-term vision for Amazon's science-based competitive, perception-preserving pricing techniques. Develop and advance price prediction models leveraging deep learning frameworks, transformer architectures, and advanced statistical methods to drive pricing accuracy at scale. Build strong collaborations. Partner with product, engineering, and science teams within Pricing & Promotions to deploy machine learning price estimation and error correction solutions at Amazon scale. Design and implement neural network-based architectures — including sequence models and transformers — for large-scale price prediction and optimization. Stay informed. Establish mechanisms to stay up to date on the latest scientific advancements in deep learning, transformer architectures, applied statistics, neural network design, probabilistic forecasting, and multi-objective optimization techniques. Identify opportunities to apply them to relevant Pricing & Promotions business problems. Keep innovating for our customers. Foster an environment that promotes rapid experimentation, continuous learning, and incremental value delivery. Leverage statistical rigor and modern deep learning approaches to validate hypotheses and drive measurable pricing improvements. Successfully execute & deliver. Apply your exceptional technical machine learning expertise — including deep neural networks, attention-based models, and applied statistical analysis — to incrementally move the needle on some of our hardest pricing problems. A day in the life We are hiring a Sr. Applied Scientist to drive our pricing optimization initiatives. We drive cross-domain and cross-system improvements through: * shape and extend our RL optimization platform - a pricing centric tool that automates the optimization of various system parameters and price inputs. * Error detection and price quality guardrails at scale. * Identifying opportunities to optimally price across systems and contexts (marketplaces, request types, event periods) Price is a highly relevant input into Stores architectures; this role creates the opportunity to drive extremely large impact (measured in Bs not Ms), but demands careful thought and clear communication. About the team The Pricing Optimization science group builds and refines Amazon's algorithmic pricing and promotion models at scale. Our team combines expertise in deep learning, transformer architectures, applied statistics, and probabilistic forecasting to develop price prediction systems that directly impact the customer experience. The team also brings hands-on experience with causal modeling and inference — including uplift modeling and treatment effect estimation — to rigorously measure the impact of pricing decisions on customer behavior and business outcomes. We partner closely with product, engineering, and business teams to take solutions from research through production deployment.
IN, KA, Bangalore
Have you ever ordered a product on Amazon and when that box with the smile arrived you wondered how it got to you so fast? Have you wondered where it came from and how much it cost Amazon to deliver it to you? If so, the WW Amazon Logistics, Business Analytics team is for you. We manage the delivery of tens of millions of products every week to Amazon’s customers, achieving on-time delivery in a cost-effective manner. We are looking for an enthusiastic, customer obsessed, Sr. Applied Scientist with good analytical skills to help manage projects and operations, implement scheduling solutions, improve metrics, and develop scalable processes and tools. The primary role of an Operations Research Scientist within Amazon is to address business challenges through building a compelling case, and using data to influence change across the organization. This individual will be given responsibility on their first day to own those business challenges and the autonomy to think strategically and make data driven decisions. Decisions and tools made in this role will have significant impact to the customer experience, as it will have a major impact on how the final phase of delivery is done at Amazon. Ideal candidates will be a high potential, strategic and analytic graduate with a PhD in (Operations Research, Statistics, Engineering, and Supply Chain) ready for challenging opportunities in the core of our world class operations space. Great candidates have a history of operations research, and the ability to use data and research to make changes. This role requires robust program management skills and research science skills in order to act on research outcomes. This individual will need to be able to work with a team, but also be comfortable making decisions independently, in what is often times an ambiguous environment. Responsibilities may include: - Develop input and assumptions based preexisting models to estimate the costs and savings opportunities associated with varying levels of network growth and operations - Creating metrics to measure business performance, identify root causes and trends, and prescribe action plans - Managing multiple projects simultaneously - Working with technology teams and product managers to develop new tools and systems to support the growth of the business - Communicating with and supporting various internal stakeholders and external audiences
US, CA, Palo Alto
About Sponsored Products and Brands The Sponsored Products and Brands team at Amazon Ads is re-imagining the advertising landscape through generative AI technologies, revolutionizing how millions of customers discover products and engage with brands across Amazon.com and beyond. We are at the forefront of re-inventing advertising experiences, bridging human creativity with artificial intelligence to transform every aspect of the advertising lifecycle from ad creation and optimization to performance analysis and customer insights. We are a passionate group of innovators dedicated to developing responsible and intelligent AI technologies that balance the needs of advertisers, enhance the shopping experience, and strengthen the marketplace. If you're energized by solving complex challenges and pushing the boundaries of what's possible with AI, join us in shaping the future of advertising. Key job responsibilities As a Machine Learning Applied Scientist, you will: * Conduct deep data analysis to derive insights to the business, and identify gaps and new opportunities * Develop scalable and effective machine-learning models and optimization strategies to solve business problems * Run regular A/B experiments, gather data, and perform statistical analysis * Work closely with software engineers to deliver end-to-end solutions into production * Improve the scalability, efficiency and automation of large-scale data analytics, model training, deployment and serving * Conduct research on new machine-learning modeling and Generative AI solutions to optimize all aspects of Sponsored Products and Brands business About the team The Ad Response Prediction team within Sponsored Products and Brands (SPB) drives personalized shopping experiences for SPB Ads across placements, pages, and devices worldwide. We achieve this through ML and GenAI solutions that include customized shopper response prediction and session-level understanding to optimize every stage of the ad-serving process, from sourcing and bidding to widget discovery and auctions. Our responsibilities include advancing response prediction through model and feature innovations and extending prediction beyond the auction stage to areas such as targeting, sourcing, and bidding.
US, NY, New York
We are seeking a Measurement and Attribution Analytics scientist to further the development and application of analytics methods to examine the complex data flows of measurement, attribution and shopper analytics and to translate deep-dives into actionable insights for our product teams. In this role you will develop new tools to analyze our advertising data to help improve the performance of our bidding algorithms, targeting and relevance systems, help advance our supply strategy, and evaluate the adoption and impact of feature releases. Key job responsibilities - Analyze data trends regarding supply, optimization, ad load, and advertising mix effects that affect advertiser performance and contribute to achieving advertiser goals. - Present papers to senior leaders on issues like feature development impact on identity recognition rates, and changes of ad selection systems to improve fill rate highlighting insights that will inform our business development and engineering roadmaps. - Identify, standardize, and operationalize KPIs to effectively measure the performance of all systems involved in ad serving, and use trend insights to inform business priorities. - Partner with engineering teams to define data logging requirements and getting these prioritized in engineering roadmaps. - Validate financial models through analysis - Develop and own ad revenue and supply intelligence analytics decks that provide ongoing deep-dives A day in the life The Measurement and Attribution Scientist will work closely with business leaders and engineers on developing common data architecture that will optimize our data logging at different grains, and will allow data interoperability from bid flow to optimization to campaign delivery. There will be an emphasis on understanding the impact of shopper journeys to and from Amazon properties and their implication on product development. The candidate will then analyze the data and present papers and ongoing reports on actionable insights. About the team The Ads Science Product Team's Mission: Work alongside those who need product data to apply objective perspective and business logic to uncover insights, advise strategic decisions, and adjust to industry changes.
IN, KA, Bangalore
Does the thought of improving one of the world’s most complex logistic systems inspire you? Is your passion to sift through hundreds of systems, processes, and data sources to solve the puzzle and identify the next big opportunity? Are you a creative big thinker who is passionate about using data to direct decision making and solve complex and large-scale challenges? Are you fascinated by the interactions between operations and strategy? Do you feel like your skills uniquely qualify you to bridge communication between teams with competing priorities? If so, then this position is for you! Come help Amazon create state-of-the-art science-driven technologies for delivering packages to the doorstep of our customers! The Last Mile Routing & Planning organization builds the software, algorithms and tools that make the “magic” of home delivery happen: our flow, sort, dispatch and routing intelligence systems are responsible for the billions of daily decisions needed to plan and execute safe, efficient and frustration-free routes for drivers around the world. Our team supports deliveries (and pickups!) for Amazon Logistics, Same Day, Amazon Grocery, Lockers, and other new initiatives across the world. Key job responsibilities In this role, your main focus will be to apply algorithms, synthesize information, identify business opportunities, provide data-driven insights and communicate business and technical requirements within the team and across stakeholder groups. You will partner closely with other scientists and engineers in a collegial environment with a clear path to business impact. We have an exciting portfolio of research areas including vehicle routing, planning for electric and autonomous vehicles, district and stops planning, ultra-fast deliveries, fleet planning, and forecasting solutions for different delivery programs leveraging the latest OR, ML, and Generative AI methods, at a global scale. Successful candidates will have a deep knowledge of Operations Research and/or Machine/Deep Learning methods, experience in applying these methods to large-scale business problems, the ability to map models into production-worthy code in Python or Java, the communication skills necessary to explain complex technical approaches to a variety of stakeholders and customers, and the excitement to take iterative approaches to tackle big research challenges.
US, WA, Bellevue
As an Applied Scientist in Amazon Fullfilment Technology, you will lead the development of agentic systems to assist with operational decision making and orchestration. You will work building full agentic systems leveraging multi-agent orchestration, tool use, memory, and action execution. You will train LLMs using a combination of rejection sampling approaches, SFT, continual post-training, and Reinforcement Learning (RL). These systems are deployed to Amazon buildings, and you will also work on rigorous offline and online evaluations. Your work will leverage the latest LLMs to develop capabilities for agentic reasoning, coding and analytics. You will also lead research projects to tackle unsolved problems, mentor interns, and author academic papers to summarize your findings for external publication. Key job responsibilities - Generating training and preference data for specific use cases (reasoning trajectories, tool traces) - Reward modeling and policy optimization for LLMs: DPO, IPO, RLHF/RLAIF with PPO/GRPO, rejection sampling. - Supervised fine-tuning on step-by-step trajectories and tool-use traces - Verbal Reinforcement Learning and Continual Learning - RL for LLMs, Offline RL and off-policy evaluation - Agentic memory/state management; episodic and semantic memory; vector search; grounding with RAG. - Evaluation: developing decision quality metrics, scaling LLM-based evaluations. About the team Amazon Fulfillment Technologies (AFT) powers Amazon's global fulfillment network. We invent and deliver software, hardware, and data science solutions that orchestrate processes, robots, machines, and people. We harmonize the physical and virtual world so Amazon customers can get what they want, when they want it. Learn more about AFT: https://tinyurl.com/AFTOverview
US, WA, Bellevue
As an Applied Scientist in Amazon Fullfilment Technology, you will lead the development of agentic systems to assist with operational decision making and orchestration. You will work building full agentic systems leveraging multi-agent orchestration, tool use, memory, and action execution. You will train LLMs using a combination of rejection sampling approaches, SFT, continual post-training, and Reinforcement Learning (RL). These systems are deployed to Amazon buildings, and you will also work on rigorous offline and online evaluations. Your work will leverage the latest LLMs to develop capabilities for agentic reasoning, coding and analytics. You will also lead research projects to tackle unsolved problems, mentor interns, and author academic papers to summarize your findings for external publication. Key job responsibilities - Generating training and preference data for specific use cases (reasoning trajectories, tool traces) - Reward modeling and policy optimization for LLMs: DPO, IPO, RLHF/RLAIF with PPO/GRPO, rejection sampling. - Supervised fine-tuning on step-by-step trajectories and tool-use traces - Verbal Reinforcement Learning and Continual Learning - RL for LLMs, Offline RL and off-policy evaluation - Agentic memory/state management; episodic and semantic memory; vector search; grounding with RAG. - Evaluation: developing decision quality metrics, scaling LLM-based evaluations. About the team Amazon Fulfillment Technologies (AFT) powers Amazon's global fulfillment network. We invent and deliver software, hardware, and data science solutions that orchestrate processes, robots, machines, and people. We harmonize the physical and virtual world so Amazon customers can get what they want, when they want it. Learn more about AFT: https://tinyurl.com/AFTOverview
US, VA, Arlington
Are you excited about making business decisions using science and data? Are you interested in supporting consumer device concepts from idea inception to launch? Do you want to work on a Science Product team focused on scaling statistics and econometrics with custom tools? If so, this may be the role for you! Amazon.com strives to be Earth's most customer-centric company. The Amazon Devices and Services team focuses on delighting customer by enabling seamless functionality in supplying, entertaining, and managing the home -- and beyond. We seek and hire the world's brightest minds, offering them a fast-paced, technologically-sophisticated, and friendly work environment, where economic theory meets real-world industry. The Decision Science team in Devices owns demand estimates and pricing recommendations of concept devices before customers know they exist. We support devices and services ranging from Echo Frames to Kindle Paperwhite to Blink Video Camera …all prior to launch. We are a cross-functional Product team working to scale Econometrics through Amazon and beyond by incorporating Science into internal facing tools and making it easier for others to do so as well. In this role, you will have input in decision meetings with Amazon senior leadership, which include go/no-go decisions for brand new devices and services and build volume decisions for manufacture prior to receiving any customer signal. You will have direct input to pricing decisions. You will leverage Science and Tools produced by the Decision Science team such as conjoint demand models to produce these recommendations. You will work with Scientists, Economists, Product Managers, and Software Developers to provide meaningful feedback about stakeholder problems to inform business solutions and increase the velocity, quality, and scope behind our recommendations. You will also have the opportunity to work on special projects to both guide the business and advance your own knowledge and understanding of specific topics. Key job responsibilities Applies expertise to develop econometric/machine learning models to measure the demand of devices and the business; Reviews models and results for other scientists, mentors junior scientists; Generates economic insights for the Devices and Services business and work with stakeholders to run the business for effectively; Describes strategic importance of vision inside and outside of team; and, Identifies business opportunities, defines the problem and how to solve it; Engages with senior scientists, business leadership outside Devices and Services to understand interplay between different business units.