Alexander Shypula holding a long-haired dachshund in a car  
My name is Alexander and I’m a researcher interested in evaluating and improving the diversity and creativity of LLMs, and in applying LLMs to complex programming tasks like program optimization and decompilation.

I’m currently a fifth year PhD student at the University of Pennsylvania, graduating in 2027, where I’m advised by Osbert Bastani. Before that, I spent a year working at MIT’s Computer Science and Artificial Intelligence Laboratory (CSAIL) with Yoon Kim (MIT CSAIL) and Jie Chen (MIT-IBM Watson AI Lab). Earlier, I was a Master’s student at Carnegie Mellon University’s (CMU) School of Computer Science in the Artificial Intelligence and Innovation program, where I was a member of Neulab and advised by Graham Neubig.

Publications

Shypula A, Bastani O, Schwartz E. “Decaf: Improving Neural Decompilation with Automatic Feedback and Search”. arXiv preprint, 2026.

Anupam S, Shypula A, Bastani O. “LLM Program Optimization via Retrieval Augmented Search”. Findings of ACL 2026.

Wong J, Orlovskiy Y, Shypula A, Luo M, Seshia S, Gonzalez J. “SimpleStrat: Diversifying Language Model Generation with Stratification”. NeurIPS 2025.

Shypula A, Madaan A, Zeng Y, Alon U, Gardner J, Hashemi M, Neubig G, Ranganathan P, Bastani O, Yazdanbakhsh A. “Automated High-Level Code Optimization for Warehouse Performance”. IEEE Micro, 2025 (Special Issue: Top Picks from the 2024 Computer Architecture Conferences). Top Picks write-up of the ICLR 2024 paper, selected as one of the 12 most significant computer-architecture papers of 2024.

Shypula A, Li S, Zhang B, Padmakumar V, Yin K, Bastani O. “Evaluating the Diversity and Quality of LLM Generated Content”. COLM 2025. An earlier version appeared at the DL4C Workshop at ICLR 2025 as “Does Instruction Tuning Reduce Diversity? A Case Study Using Code Generation.”

Shypula A*, Madaan A*, Zeng Y, Alon U, Gardner J, Yang Y, Hashemi M, Neubig G, Ranganathan P, Bastani O, Yazdanbakhsh A. “Learning Performance-Improving Code Edits”. ICLR 2024 (Spotlight); selected for IEEE Micro Top Picks 2025. * Equal contribution. The approach was deployed in production at Google as ECO, whose authors report savings equivalent to over 500k normalized CPU cores per quarter.

Tang Z, Agarwal M, Shypula A, Wang B, Wijaya D, Chen J, Kim Y. “Explain-then-Translate: An Analysis on Improving Program Translation with Self-generated Explanations”. EMNLP 2023.

Shypula A, Yin P, Lacomis J, Le Goues C, Schwartz E, Neubig G. “Learning to Superoptimize Real-world Programs”. Best Paper, inaugural Deep Learning for Code (DL4C) Workshop at ICLR 2022.

Full list on Google Scholar.

Contact

You can reach me at shypula 👨‍💻 seas ☕ upenn 🚴 edu.

Background

I’ve been interested in improving the performance of deep learning models for code and symbolic domains since the first systems course I took. AI in these symbolic domains is interesting because of the feedback loops available: code can be analyzed and executed in ways natural language cannot. Because of this, when I started my PhD I hoped research would focus not only on imitating human programmers, but on how AI models can teach us to program and reason better, like AlphaGo’s move 37. We’re seeing this today with the incredible capacity of tools like Claude Code, Codex, and LLM theorem provers.

On the programming side, I’ve worked on program optimization, where a model has to find faster code that is still correct, and on decompilation, where it has to recover readable source from compiled binaries. Both are hard, verifiable tasks where feedback from running the code, rather than imitation alone, is what lets the model improve.

I think a major vector for self-improving AI is going to be the loop of generating synthetic data, filtering it, and then re-training on it. My work on the diversity and creativity of LLMs has two purposes. One is avoiding the “artificialness” of mode collapse and the syntactic slop we get from frontier LLMs. The other is understanding how diversity lets models generate useful synthetic data, or, in RL language, “explore” novel problems and self-improve.

My family roots are in Poland: my parents fled an oppressive authoritarian regime, and their parents lived through genocide and the uprooting of mass migration after the Second World War. We’ve been fortunate and unfortunate in different ways, and I’ve learned that the purpose of life lies in overcoming obstacles, not in outcomes.