Alexander Long is seen wearing a suit, speaking at podium, the banner behind him and to the right says data to decisions CRC — Alexander Long, an applied scientist in Australia, said he initially was set to follow his father's career path in the oil and gas industry — until he discovered reinforcement learning.

Computer vision

How a passion for reinforcement learning guided Alexander Long’s trajectory

The field motivated him to pursue a PhD, which eventually led him to Amazon.

June 24, 2022

6 min read

Alexander Long had his mind set on working in the oil and gas industry, following in his father’s footsteps. The sector is a big employer of electrical engineers in his home country of Australia, so it was a natural path after getting his bachelor’s degree at The University of Queensland (UQ).

In 2013, as Long was preparing to graduate, he became the first student selected for a collaboration between UQ and the Technical University of Munich (TUM). He spent two years in Germany, completing simultaneous master’s degrees in electrical engineering — both at UQ and at TUM. That’s when he heard about reinforcement learning (RL) for the first time — and he quickly realized he wanted to go deeper.

“Reinforcement learning is one way to frame the problem of making optimal actions,” Long explained. “Chess is a good example of a situation where you have an objective — winning the game — and you have to take a bunch of sequential steps to meet that objective. But you don’t get any concrete feedback until after you’ve made 20 or 30 moves.” The same framework can be used to solve a multitude of problems, from winning a game to optimizing a refinery or controlling a nuclear fusion reactor.

The widespread applications for reinforcement learning fascinated Long. But, he notes, the method has some significant drawbacks. “One of those is you need huge amounts of interactions with an environment before you can learn how to act well,” he explained.

Learning faster

See Amazon's Australia research locations

After completing his master’s program, Long pursued a PhD in computer science at the University of New South Wales (UNSW). He wanted to explore the challenge of how to help RL models become more data efficient by learning from fewer interactions.

The outcome was “Fast and Data Efficient Reinforcement Learning from Pixels via Non-Parametric Value Approximation”, a paper that was presented as part of an AAAI 2022 poster session.

It was very surprising; the algorithm was on par with all the best methods in terms of data efficiency, but it was about 100 times faster in terms of computation time.

Alexander Long

The paper notes that previous advances in RL algorithm efficiency “have been achieved at the cost of increased sample, and computational complexity.” That added complexity “presents a major roadblock” for online, real-world settings. In their paper, the researchers presented “Nonparametric Approximation of Inter-Trace returns (NAIT), an algorithm that is both computation and sample efficient.”

“I was poking around that area, doing baseline work, and I found there was a very basic method that could be modernized by adding a couple of innovations, but nothing crazy, and that it worked extremely well,” he says. “It was very surprising; the algorithm was on par with all the best methods in terms of data efficiency, but it was about 100 times faster in terms of computation time.”

This screenshot shows the top third of the Amazon.com homepage as of Jan. 19, 2022

Joining Amazon

When Long saw that Amazon was opening an office in Australia in 2021, he focused his energies on getting a job there. He did that by contacting his future boss, Anton van den Hengel, director of applied science at Amazon.

“I emailed him three times, pestering him for a job,” he recalled. Eventually he gained an interview for an internship. His first interview didn’t lead to a role, but his second did.

Anton van den Hengel is seen smiling into the camera, with some office buildings in the background

Start-up mindset

At the end of his internship, Long went through a set of interviews and presented the work he had done over that period to help secure a full-time position as an applied scientist. Van den Hengel said the decision to hire Long was easy. “He has great skills, and a strong publication record. More than that though, he demonstrated the ability to apply and extend the state of the art in ML research. That’s what we’re seeking.”

I was told to set my own direction, work at my own pace, and let’s see what you do at the end of six months. The other exceptional thing about the internship was hanging out with some of the smartest people.

Alexander Long

Looking back on his internship, Long said his startup experience led him to assume a big company like Amazon meant he wouldn’t have as much freedom and would be told exactly what to do.

“It was not like that at all,” he noted. “I was told to set my own direction, work at my own pace, and let’s see what you do at the end of six months.”

“The other exceptional thing about the internship was hanging out with some of the smartest people,” Long said. In his first weeks as an intern, he was in the process of getting his PhD paper published and shared a draft with one of his colleagues, who quickly suggested invaluable changes. “He knew all these little things that no one at my university knew. And you have interactions like that all the time.”

Long compares his experience at Amazon with that of his father’s in oil and gas, where small improvements in efficiency could have tens or hundreds of millions of dollars of business impact. “It’s awesome that one person or a group of people can sit down, think hard, and have a disproportionate effect on both customers and the business. There are very few places where that can occur.”

See science jobs in Australia

Amazon has openings for data scientists, applied scientists, machine learning scientists, and more at Amazon's offices in Australia.

Browse open roles

About the Author

Mariana Lenharo

Mariana is a science and health journalist based in Brazil.

How a passion for reinforcement learning guided Alexander Long’s trajectory

The field motivated him to pursue a PhD, which eventually led him to Amazon.

Learning faster

Joining Amazon

Start-up mindset

Related content

Work with us