Papers by Adriano Koshiyama

5 papers
From Text to Emoji: How PEFT-Driven Personality Manipulation Unleashes the Emoji Potential in LLMs (2025.findings-naacl)

Copied to clipboard

Challenge: Methods like prompt-based In-Context Knowledge Editing and gradient-based Model Editor Networks (MEND) show irregularity and variability; IKE depends on the prompt, leading to variability and sensitivity; MEND yields inconsistent and gibberish outputs.
Approach: They employ Opinion QA Based Parameter-Efficient Fine-Tuning (PEFT) to manipulate the Big Five personality traits: Openness, Conscientiousness, Extraversion, Agreeableness, and Neuroticism.
Outcome: The proposed methods show that they are more accurate than prompt-based IKE and gradient-based MEND outputs.
HyPA-RAG: A Hybrid Parameter Adaptive Retrieval-Augmented Generation System for AI Legal and Policy Applications (2025.naacl-industry)

Copied to clipboard

Challenge: Large Language Models (LLMs) face limitations due to outdated knowledge, hallucinations, and poor reasoning in complex contexts.
Approach: They propose a Hybrid Parameter-Adaptive RAG system for the AI legal domain with NYC Local Law 144 as the test case.
Outcome: The proposed system improves retrieval accuracy, response fidelity, and contextual precision on NYC Local Law 144 . Empirical evidence indicates that many AI tools overstate their ability to prevent hallucinations in legal and policy contexts.
LibVulnWatch: A Deep Assessment Agent System and Leaderboard for Uncovering Hidden Vulnerabilities in Open-Source AI Libraries (2025.acl-srw)

Copied to clipboard

Challenge: Open-source AI libraries present significant, underexamined risks spanning security, licensing, maintenance, supply chain integrity, and regulatory compliance.
Approach: They propose a system that leverages large language models and agentic workflows to perform deep, evidence-based evaluations of open-source AI libraries.
Outcome: The proposed system covers up to 88% of OpenSSF Scorecard checks and uncovers 19 additional risks per library.
SAGED: A Holistic Bias-Benchmarking Pipeline for Language Models with Customisable Fairness Calibration (2025.coling-main)

Copied to clipboard

Challenge: Existing benchmarks for large language models fail to detect bias due to limited scope, contamination, and lack of a fairness baseline.
Approach: They propose a benchmarking pipeline to detect biases in large language models . they use metrics for max disparity, impact ratio, and bias concentration to analyze disparity .
Outcome: SAGED(bias) is the first holistic benchmarking pipeline to address biases in large language models.
JobFair: A Framework for Benchmarking Gender Hiring Bias in Large Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: a framework for benchmarking hierarchical gender hiring bias in Large Language Models (LLMs) is developed to protect vulnerable demographic groups.
Approach: They propose a framework for benchmarking hierarchical gender hiring bias in Large Language Models for resume scoring.
Outcome: The proposed framework reveals significant issues of reverse gender hiring bias and overdebiasing in ten state-of-the-art LLMs.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations