About this CFP
Amazon's Devices and Services organization delivers a constellation of devices, services, and experiences that empower customers to stay connected, safe, informed, productive, educated, and entertained. The portfolio includes Echo, Alexa+, Fire TV, Fire tablets, Kindle, Ring, Blink, and eero. Our science and engineering teams work across audio, camera, and sensor systems, and on the on-device intelligence that turns what a device measures into something useful to the customer.
Compact, battery-powered devices sense the world through small apertures, modest optics, and low-power silicon, continuously and within a fixed power and thermal budget. What such a sensor produces is dim, noisy, and geometrically imperfect, and the capture itself may be incomplete: the device was in the wrong place, a moment was missed, or recording was sparse. It still has to become something useful: a clear image or video, a compact representation a model can consume, a reliable estimate of where things are, or the visual experience the user intended. This CFP solicits research that closes the gap between what a compact device captures and what the user needs, whether through learned on-device processing or generative methods that work from limited observations.
We welcome proposals in the following research topics:
Research topics
Proposals may address more than one topic.
1. Generative Visual Experience Beyond Capture Constraints. Compact devices that capture continuously or opportunistically produce imagery that is limited not only by sensor and environmental conditions but by the circumstances of the capture itself: the device was too far from the subject, the viewing angle was poor, a key moment was missed or only partially recorded, or the capture was intermittent and sparse rather than continuous.
These are content-level limitations that image restoration cannot address. This research topic solicits work on generating the visual experience the user intended from the observations the device actually made. Of Interest:
- Recovering or synthesizing key visual moments from incomplete, poorly timed, or partially occluded capture
- Generating plausible views from perspectives the device did not occupy, grounded in the observations that were actually made
- Producing temporally coherent, continuous visual experiences from sparse or discontinuous input
Evaluation methodology is part of the challenge: proposals should describe how they would measure success, particularly where standard benchmarks do not exist.
2. Efficient visual encoding for on-device vision-language models. Determining the best way to encode visual input for a downstream vision-language model under a fixed power and bandwidth budget. Open questions include learned compression versus semantic tokenization, single-stage versus distributed encoding, and how to evaluate the result against bits spent, image degradation, and task success together rather than reconstruction quality alone. The difficulty is that the available approaches each trade away one of three properties: cost, relevance-awareness, and open-endedness.
- Pixel-level codecs are cheap and general, but spend budget uniformly regardless of what matters
- Semantic tokenizers know what matters, but are typically too costly to run continuously
- Narrow task-specific encoders are cheap and selective, but do not generalize to open-ended queries
We are interested in approaches that improve on these trade-offs.
3. Learned multimodal sensor fusion for object tracking. Estimating the position and motion of objects from raw inertial, visual, and ranging streams, which are asynchronous and multi-rate, by learning a shared latent state rather than authoring a measurement and dynamics model by hand. Each sensor is unreliable in a different way, and some limits of the hand-authored approach are structural rather than a matter of tuning. Of interest:
- Combining the raw streams without losing cross-sensor structure, and supervising training without ground-truth position, which is unavailable at scale (for example self-supervised early fusion)
- How much of the classical measurement and dynamics model a learned system can absorb, and how much still has to be specified by hand
- Calibration-free deployment, with mounting offset and sensor bias inferred rather than measured per unit
- Continued estimation through extended sensor occlusion, where inertial data alone drifts quickly
Theoretical advances, creative new ideas, and practical applications are all welcome.
Timeline
Submission period: October 1 — November 4, 2026 (11:59PM Pacific Time).
Decision letters will be sent out in February 2027.
Award details
Selected Principal Investigators (PIs) may receive the following:
- Unrestricted funds, no more than $80,000 USD on average
- AWS Promotional Credits, no more than $40,000 USD on average
- Training resources, including AWS tutorials and hands-on sessions with Amazon scientists and engineers
Awards are structured as one-time unrestricted gifts. The budget should include a list of expected costs specified in USD, and should not include administrative overhead costs. The final award amount will be determined by the awards panel.
Eligibility requirements
Please refer to the ARA Program rules on the Rules and Eligibility page.
Proposal requirements
Proposals should be prepared according to the ARA proposal template and are encouraged to be a maximum of 4 pages, not including Appendices. In addition, to submit a proposal for this CFP, please also include the following information:
- Which of the research topics above your proposal targets
- How your approach differs from or builds upon existing methods in the area
- The constraints the method operates under and evidence it can meet them
- How you will evaluate the work, including the trade-off being optimized and the metrics used
- What will be achieved within the one-year award period
- Please list the open-source code, datasets, or benchmarks you plan to contribute
- Please list the AWS tools and services you will use
Selection criteria
Proposals will be reviewed by a panel of Amazon scientists and engineers. Proposals will be evaluated on the following:
- Creativity and quality of the scientific content
- Fit to one or more of the research topics above
- Credibility under the constraints relevant to the targeted research topic
- Feasibility of the proposed approach within the one-year award period
- Interest expressed in open-sourcing code, datasets, and benchmarks
Expectations from recipients
To the extent deemed reasonable, award recipients may acknowledge support from ARA (e.g., (“Research reported in this [publication/press release] was supported by an Amazon Research Award, [Cycle /Year].“). Award recipients will inform ARA of publications, presentations, code and data releases, blogs/social media posts, and other speaking engagements referencing the results of the supported research or the Award. Award recipients are expected to provide updates and feedback to ARA via surveys or reports on the status of their research. Award recipients will have an opportunity to work with ARA on an informational statement about the awarded project that may be used to generate visibility for their institutions and ARA.