Papers by Henryk Michalewski

6 papers
A Simple, Yet Effective Approach to Finding Biases in Code Generation (2023.findings-acl)

Copied to clipboard

Challenge: Recent work shows that large language models can generate code on par with humans . however, data-driven approaches may not be sufficient for acquiring reasoning skills .
Approach: They propose a framework that automatically identifies subtle cues a code generation model might exploit . they propose an automated intervention mechanism reminiscent of adversarial testing .
Outcome: The proposed framework can be used as a data transformation technique during fine-tuning, acting as reversal strategy.
Hierarchical Transformers Are More Efficient Language Models (2022.findings-naacl)

Copied to clipboard

Challenge: Transformers are impressive but inefficient and costly, which limits their applications and accessibility.
Approach: They first use different ways to downsample and upsamplify activations in Transformers to make them hierarchical.
Outcome: The proposed model outperforms Transformers on the ImageNet32 and enwik8 benchmarks.
Towards Robust Mathematical Reasoning (2025.emnlp-main)

Copied to clipboard

Challenge: IMO-Bench is a suite of advanced reasoning benchmarks that targets the international mathematical Olympiad level.
Approach: They propose IMO-Bench, a suite of advanced reasoning benchmarks that targets the level of the international mathematical Olympiad.
Outcome: IMO-Bench is a suite of advanced reasoning benchmarks that targets the level of the international mathematical Olympiad.
Beyond Lines and Circles: Unveiling the Geometric Reasoning Gap in Large Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in Large Language Models (LLMs) demonstrate increasing proficiency in complex mathematical and algorithmic tasks, yet their geometric reasoning skills are underexplored.
Approach: They propose a framework that enhances LLMs’ reasoning potential through a multi-agent system conducting internal dialogue.
Outcome: The proposed framework enhances LLMs’ reasoning potential through a multi-agent system conducting internal dialogue.
Natural Language to Code Generation in Interactive Data Science Notebooks (2023.acl-long)

Copied to clipboard

Challenge: Data scientists use computational notebooks to perform data wrangling and analytic tasks.
Approach: They build a benchmark program that synthesizes programs given NL intents from users by using a Python code language model.
Outcome: The proposed model outperforms public code LMs in a dataset of 1078 code generation problems using the pandas data analysis framework in data science notebooks.
Measuring and Improving BERT’s Mathematical Abilities by Predicting the Order of Reasoning. (2021.acl-short)

Copied to clipboard

Challenge: a common language model for word math problems lacks mathematical abilities . a data-driven approach to solving word problems is lacking in many areas .
Approach: They propose to train a language model with mathematical abilities to teach word maths . they propose to use semi-formal steps to explain how math results are derived .
Outcome: The proposed model achieves better outcomes than baseline models and on-par with more tailored models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations