What is Build on Trainium?
Build on Trainium is a $110MM credit program focused on AI research and university education to support the next generation of innovation and development on AWS Trainium. Amazon Trainium chips are purpose-built for high-performance deep learning (DL) training of generative AI models, including large language models (LLMs) and latent diffusion models. Build on Trainium provides compute credits to novel AI research on Trainium, investing in leading academic teams to build innovations in critical areas including new model architectures, ML libraries, optimizations, large-scale distributed systems, and more. This multi-year initiative lays the foundation for the future of AI by inspiring the academic community to utilize, invest in, and contribute to the open-source community around Trainium. Combining these benefits with the Neuron software development kit (SDK) and the Neuron Kernel Interface (NKI), AI researchers can innovate at scale in the cloud.
What are AWS Trainium and Neuron?
Amazon Trainium is an AI chip developed by Amazon for accelerating building and deploying machine learning models. Built on a specialized architecture designed for deep learning, Trainium accelerates the training and inference of complex models with high output and scalability, making it ideal for academic researchers looking to optimize performance and costs. This architecture also emphasizes sustainability through energy-efficient design, reducing environmental impact. Amazon has established a dedicated Trainium research cluster featuring up to 40,000 Trainium chips, accessible via Amazon EC2 Trn2 instances. These instances are connected through a non-blocking, petabit-scale network using Amazon EC2 UltraClusters, enabling seamless high-performance ML training. The Trn2 instance family is optimized to deliver substantial compute power for cutting-edge AI research and development. This unique offering not only enhances the efficiency and affordability of model training but also presents academic researchers with opportunities to publish new papers on underrepresented compute architectures, thus advancing the field.
Focus on Agentic, recursive improvement of kernels and frameworks
Getting the most out of an AI accelerator means optimizing the full stack, from frameworks and serving systems down to individual kernels. This is iterative and expert work. Generative AI can change this through the agent loop: an agent generates, runs, profiles, and repairs its work against hard evidence from the compiler, profiler, and correctness checks, then proposes the next candidate without a human in every cycle. The more compelling opportunity is recursive self-improvement, where a system gets better at the task over time by refining its own models, context, tools, and knowledge base.
We seek proposals advancing agentic, self-improving frameworks for kernel and end-to-end optimization, across the following areas. The strongest will connect the generation, evaluation, and learning loops rather than treating them in isolation.
1. Generative-AI-Guided generation and optimization, from kernels to end-to-end workloads
This area covers agents that generate and optimize code from a single kernel up to a whole training or inference system. We seek proposals across the following:
Kernel generation, optimization and auto-tuning
Agentic systems that generate correct kernel code from natural-language specifications, other kernel languages, or tensor-algebra formulations like einsum. These systems should ground their output in existing corpora, documentation, and Instruction Set Architecture (ISA) constraints, then run a compile-execute-repair loop until the result is verified. We also seek agents that optimize correct kernels via instruction scheduling, memory-access restructuring (tiling, double buffering), and auto-tuning of tile sizes, loop order, fusion, and allocation against the target memory hierarchy. These techniques are illustrative, not exhaustive, and novel strategies are welcome. Of particular interest are agents that keep an LLM in the loop, reasoning over profiler feedback to propose the next candidate.
System and workload optimization
A kernel that is fast on its own does not guarantee a fast system; the largest gains often come from how kernels are combined, scheduled, and fed with data. We seek agentic systems that reason over and reshape the whole framework. On the inference side, that means the serving machinery itself: scheduling prefill and decode, laying out and paging the KV cache, routing requests across devices, and overlapping communication with computation under a chosen sharding configuration, with the system finding the best sharding strategy and tuning the corresponding kernels. On the training side, the same extends to sharding, collective placement, pipeline scheduling, and memory planning across the model graph. We especially encourage agents that can optimize to a target metric that practitioners care about.
Efficient optimization loops
Whole model optimization is expensive in agent tokens and tool time, so the loop itself must be efficient. We seek agents, or coordinated systems of agents, that can quickly identify which candidate optimizations help and which regress. They should distinguish standalone changes from those whose effects cascade and require co-optimization. They should rank candidates by expected impact relative to their cost in iterations, time, and tokens, focusing on budget where it matters most. The goal is to compress whole-model optimization from hours to minutes. We especially welcome order-of-magnitude reductions in both time and cost.
2. Performance debugging, profiling, and explainability
AI-assisted tools for diagnosing performance and explaining it in terms a human can act on. This includes localizing bottlenecks from profiler traces, such as whether a kernel is bound by compute, memory, or data movement. It includes natural language explanations of why a kernel underperforms, ranked fix suggestions, and root cause analysis for regressions. We are also interested in how agents use the profiler. Computer-use agents could drive existing tools built for humans, and others could reimagine the profiler itself with agents as the primary consumer.
3. Correctness, integrity, and evaluation criteria
Generative AI that generates and optimizes kernels introduces a distinct failure mode: outputs that appear to "pass" but do not represent genuine optimization. This area covers both the detection of such failures and the standards used to judge success.
Reward hacking and exploit detection
Methods for detecting kernels that exploit gaps in the test harness. Examples include hardcoded outputs, code that only works on the test tensors or merely copies the reference corpora, or speedups won by disabling correctness paths, skipping edge cases, or dropping numerical precision. Relevant techniques include static and dynamic analysis to flag suspicious patterns, and adversarial or randomized tests that stress kernels beyond the evaluation distribution.
Accuracy benchmarks and acceptance criteria
Standards for judging whether a kernel is genuinely correct and fast, from functional equivalence and accuracy tolerances to rubrics against human authored baselines. Of particular interest are datasets and benchmarks for how well LLMs write kernels, built from curated (specification, reference kernel) pairs, that can serve as shared ground truth. We also welcome criteria for when a faster but slightly less accurate kernel is acceptable, and how that tradeoff should be certified and disclosed.
4. SFT and RL fine-tuning of open-source models for kernel and system optimization
Kernel and system optimization is a specialized skill, and open models taught to master it can be run in a recursive loop and reproduced, extended, and built on by the community. We seek proposals that adapt open-source foundation and code models into specialized optimization agents through supervised fine-tuning (SFT) and reinforcement learning (RL), including:
Fine-tuning objectives and long horizon training
This includes SFT on curated corpora, documentation, and paired examples (specification to kernel, or unoptimized to optimized), and RL that turns compiler feedback, profiler metrics, or correctness checks into reward signals. A key challenge is the long-horizon rollout, where an agent generates, compiles, runs, profiles, and repairs before any reward arrives. We seek methods for credit assignment across these long, sparsely rewarded trajectories.
Reward design and task specific evaluation
Reward design is central here, and we seek signals that capture genuine quality. Distinct tasks demand distinct rewards and evaluations. Optimizing a single kernel is judged on that kernel's own correctness and performance, whereas optimizing an end-to-end model is judged on system-level metrics over a full run, and training and inference differ again. A reward that fits one can mislead another. We invite designs that reflect these differences, and we encourage teams to release their models and training data openly.
Learning beyond model weights
Weights are not the only thing that can be learned. We are equally interested in using RL, alongside or instead of SFT, to improve the agent's scaffolding: the context it retrieves, the tools it uses, and the knowledge base it draws on. Updating what an agent knows and does, rather than only its weights, is easier to inspect, audit, and share.
Timeline
Submission period: October 1 — November 4, 2026 (11:59PM Pacific Time).
Decision letters will be sent out in February 2027.
Award details
Selected Principal Investigators (PIs) may receive the following. Applicants are encouraged to request AWS Promotional Credits in one of two ranges:
- AWS Promotional Credits, up to $50,000
- AWS Promotional Credits, up to $250,000 and beyond
- AWS Trainium training resources, including AWS tutorials and hands-on sessions with Amazon scientists and engineers
Awards are structured as one-time unrestricted gifts. The budget should include a list of expected costs specified in USD, and should not include administrative overhead costs. The final award amount will be determined by the awards panel.
Your receipt and use of AWS Promotional Credits is governed by the AWS Promotional Credit Terms and Conditions, which may be updated by AWS from time to time.
Eligibility requirements
Please refer to the ARA Program rules on the Rules and Eligibility page.
Proposal requirements
Proposals should be prepared according to the proposal template.
PIs are encouraged to exemplify how their proposed techniques or research studies advance kernel optimization, agentic and self-improving systems, LLM innovation, distributed systems, or developer efficiency. PIs should either include plans for open source contributions or state that they do not plan to make any open source contributions (data or code) under the proposed effort. Proposals for this CFP should be prepared according to the proposal template and are encouraged to be a maximum of 4 pages, not including Appendices.
Selection criteria
Proposals will be evaluated on the following:
- Creativity and quality of the scientific content
- Potential impact to the research community and society at large
- Interest expressed in open-sourcing model artifacts, datasets, and development frameworks
- Intention to use and explore novel hardware for AI/ML, primarily AWS Trainium and Inferentia
Expectations from recipients
To the extent deemed reasonable, award recipients may acknowledge support from ARA (e.g., (“Research reported in this [publication/press release] was supported by an Amazon Research Award, [Cycle /Year].“). Award recipients will inform ARA of publications, presentations, code and data releases, blogs/social media posts, and other speaking engagements referencing the results of the supported research or the Award. Award recipients are expected to provide updates and feedback to ARA via surveys or reports on the status of their research. Award recipients will have an opportunity to work with ARA on an informational statement about the awarded project that may be used to generate visibility for their institutions and ARA.