Conversational AI

"I want machines to write as fluently as humans"

Amazon Machine Learning Fellow Jiao Sun works on strategies to control text generation.

December 13, 2022

5 min read

What if artificial intelligence could help an aspiring author write a novel? Or coach people to improve the quality of their writing? Could machines learn how to make jokes? Inspired by these questions, computer scientist Jiao Sun has been exploring the potential of AI-generated text as a PhD candidate at the University of Southern California (USC).

After a four-month internship at Alexa AI last spring, she is now starting her journey as an Amazon Machine Learning Fellow for the 2022–23 academic year and hopes to continue developing text-generation models that enhance the interaction between humans and AI.

Jiao Sun is seen standing next to some posters she presented at EMNLP 2022 — Jiao Sun has been exploring the potential of AI-generated text as a PhD candidate at the University of Southern California. After an internship at Alexa AI last spring, she is starting her journey as an Amazon Machine Learning Fellow for the 2022–23 academic year.

While Sun is passionate about the potential of natural language generation, she also believes it’s important to develop tools that improve human control over machine-created content. She is also cautiously optimistic about the surge in popularity surrounding text generation models.

“I am thrilled to see more and more great models in the space of text generation in recent years,” she says. “It can help spur more innovation for text generation field, but might also obsolete some research, and even some research directions. Personally, my research philosophy is to work on research that is agnostic to model choices and creative by itself.”

One of her research goals is to improve the quality, fairness, and reliability of that content to achieve what she calls trustworthy text generation.

For example, she and her colleagues recently investigated the presence of gender stereotypes in greeting card messages written by both humans and machines. The research — which received the Best Paper Honorable Mention in the 2022 CHI Conference on Human Factors in Computing Systems, an international conference on human-computer interaction — led to the development of a writing assistant tool to combat those biases.

Protecting authors’ privacy

Sun is still in the early stages of her fellowship, but one area of research she would like to explore during the program is using AI to ensure author privacy, which she sees as another aspect of trustworthy text generation.

She notes that natural language processing techniques can be used to infer the authorship of articles and documents based on the author’s writing style, especially if the author has multiple articles published online.

But what if, for some reason, the author wants to remain anonymous?

“We're thinking about ways we can rewrite something in a way that maintains the semantics from your text while keeping the authorship protected,” Sun says. The idea is to develop AI models that rephrase contents to remove stylistic fingerprints that could give away who the author is.

jiao sun emnlp.png — Thanks to an Amazon travel grant, Jiao Sun was able to present her research in person at the recent EMNLP 2022 conference in Abu Dhabi. “This grant gave me the opportunity of traveling to what was my first in-person conference in my entire PhD,” she says.

During the program, Sun is being mentored by Qian Hu, an applied scientist at Amazon Alexa AI, with whom she connects regularly to discuss her research.

“That is not only helpful for my career, but just having this connection with another smart person helps me shape my research in the right direction,” she says.

The Amazon Machine Learning Fellowship is a program offered annually to doctoral students by the USC + Amazon Center on Secure and Trusted Machine Learning, a joint research center focused on the development of new approaches to ML privacy, security, and trustworthiness. In addition to Sun, Sina Shaham and Yunhao Ge are also ML Fellows this academic year.

‘What did the sushi say to the bee?’

During her internship at Amazon last spring, Sun worked with Amazon scientists Alessandra Cervone, Anjali Narayan-Chen, Tagyoung Chung, Shuyang Gao, Jing Huang, Yang Liu, Shereen Oraby, and Amazon Visiting Academic Violet Peng on two papers that were accepted at the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP).

Are you excited about pun generation? In #EMNLP2022, we have two works accepted in the main conference:

1️⃣ Context-Situated Pun Generation 👉 a brand-new task!
2️⃣ ExPUNations: Augmenting Puns with Keywords and Explanations 👉 a new dataset!

Learn more! 🧵👇
— Jiao Sun ✈️ EMNLP 2022 (@sunjiao123sun_) December 8, 2022

Both works are done during my internship @AmazonScience with awesome @VioletNPeng @anjalisaa @shrnrby @Ale_Cervone @iuaaui Yang Liu, Tagyoung Chung and Jing Huang!
— Jiao Sun ✈️ EMNLP 2022 (@sunjiao123sun_) December 8, 2022

“During my internship, they gave me a lot of really precious feedback. And they have continued to support me, even after my internship ended.”

Both papers explore the challenging task of explaining humor to machines. Sun notes that we often take for granted the knowledge required to understand simple puns. But imagine having to explain a play on words to a non-native speaker or a small child.

“For machines to understand jokes, they need to learn from a huge knowledge base,” she says.

Sun and her coauthors first developed a dataset of pun keywords and explanations, which was appropriately named ExPUNations. She worked on an existing dataset of puns, asking annotators to evaluate whether a given text was intended to be a joke, how funny it was to them, and what about it was funny.

Take the joke: “What did the sushi say to the bee? ‘Wasabi.’” “If I were the annotator, I would say this is funny because wasabi sounds like ‘What’s up, bee?’ That's the funniness of it,” Sun says. The annotators were also asked to select the keywords of the pun. In this case, those would be “sushi,” “bee,” and “wasabi.”

"I want machines to write as fluently as humans"

Amazon Machine Learning Fellow Jiao Sun works on strategies to control text generation.

Protecting authors’ privacy

‘What did the sushi say to the bee?’

Related content

Work with us