Papers by Tiago Almeida
Conceptual Hierarchies within LLMs (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing literature has explored abstraction within large language models (LLMs). |
| Approach: | They generate a dataset of semantic hierarchies and investigate their storage locations in six LLMs using activation patching, a causal intervention technique. |
| Outcome: | The results show that concepts at finer levels of granularity are stored around 61-78% of the time before those at coarser levels. |
Benchmarking a transformer-FREE model for ad-hoc retrieval (2021.eacl-main)
Copied to clipboard
| Challenge: | a recent study compares transformer-based models with a greener and more sustainable alternative. |
| Approach: | They compare transformer-based models with a "greener and more sustainable" alternative . they show that transformer-like models can be used in real-world retrieval applications . |
| Outcome: | The lighter model achieves a speedup of 20 times in training and 7 to 47 times in inference while maintaining a comparable retrieval performance. |
Dense Template Retrieval for Customer Support (2022.coling-1)
Copied to clipboard
| Challenge: | Templated answers are used to cover a wide range of topics, but the number of templates is often too high for an agent to manually search. |
| Approach: | They propose a dense retrieval framework that adapts a standard in-batch negatives technique to support unpaired sampling of queries and templates. |
| Outcome: | The proposed approach improves performance and training speed over more standard methods. |
Exploring efficient zero-shot synthetic dataset generation for Information Retrieval (2024.findings-eacl)
Copied to clipboard
| Challenge: | Recent advances in large language models offer a new avenue of generating synthetic training data to train neural retrieval models for unlabelled data collections. |
| Approach: | They propose a method to generate high-quality synthetic datasets using a small language model and a filtering mechanism to ensure the quality of generated questions. |
| Outcome: | The proposed method outperforms unsupervised retrieval methods such as BM25 and pretrained monoT5. |
A Framework for Fine-Grained Complexity Control in Health Answer Generation (2025.acl-srw)
Copied to clipboard
| Challenge: | Health literacy is the ability to obtain, process, and understand basic health information. |
| Approach: | They propose a framework for automatically generating health answers at multiple, precisely controlled complexity levels. |
| Outcome: | The proposed framework allows users to generate health questions at multiple complexity levels. |