TY - GEN
T1 - Can Typos Cause Harm? The Impact of Imperfect Input on LLM Safety
AU - Zinjad, Saurabh
AU - Bhattacharjee, Amrita
AU - Beigi, Alimohammad
AU - Liu, Huan
N1 - Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Switzerland AG 2026.
PY - 2026
Y1 - 2026
N2 - Large Language Models (LLMs) are increasingly being used in sensitive domains such as healthcare and education, where safety is critical. While much LLM safety research focuses on deliberate attacks (e.g., jailbreaks, prompt injections), the impact of benign but imperfect user input, such as typos or paraphrasing remains underexplored. In this study, we investigate how such semantic-preserving perturbations affect the safety behavior of aligned LLMs. We find that these small perturbations can increase instability, e.g., causing unsafe responses to flip into safe refusals. This bidirectional instability reveals that current safety alignment mechanisms are fragile and context-dependent. Using a wide set of realistic perturbations in harmful queries, we evaluated several open-source models. Our results highlight the need for more context-aware and model-sensitive evaluation frameworks and training methods that ensure robust behaviors in the face of natural and noisy input. Github: https://github.com/Ztrimus/llm-sensitivity.
AB - Large Language Models (LLMs) are increasingly being used in sensitive domains such as healthcare and education, where safety is critical. While much LLM safety research focuses on deliberate attacks (e.g., jailbreaks, prompt injections), the impact of benign but imperfect user input, such as typos or paraphrasing remains underexplored. In this study, we investigate how such semantic-preserving perturbations affect the safety behavior of aligned LLMs. We find that these small perturbations can increase instability, e.g., causing unsafe responses to flip into safe refusals. This bidirectional instability reveals that current safety alignment mechanisms are fragile and context-dependent. Using a wide set of realistic perturbations in harmful queries, we evaluated several open-source models. Our results highlight the need for more context-aware and model-sensitive evaluation frameworks and training methods that ensure robust behaviors in the face of natural and noisy input. Github: https://github.com/Ztrimus/llm-sensitivity.
KW - Input Sensitivity
KW - Prompt Perturbations
KW - Robustness Evaluation
KW - Safety Alignment
KW - Safety Flip Rates
UR - https://www.scopus.com/pages/publications/105020013213
UR - https://www.scopus.com/pages/publications/105020013213#tab=citedBy
U2 - 10.1007/978-3-032-07715-8_23
DO - 10.1007/978-3-032-07715-8_23
M3 - Conference contribution
AN - SCOPUS:105020013213
SN - 9783032077141
T3 - Lecture Notes in Computer Science
SP - 233
EP - 243
BT - Social, Cultural, and Behavioral Modeling - 18th International Conference, SBP-BRiMS 2025, Proceedings
A2 - Thomson, Robert
A2 - Pyke, Aryn A
A2 - Renshaw, Scott
A2 - Park, Patrick
A2 - Al-khateeb, Samer
A2 - Burger, Annetta
PB - Springer Science and Business Media Deutschland GmbH
T2 - 18th International Conference on Social Computing, Behavioral-Cultural Modeling and Prediction and Behavior Representation in Modeling and Simulation, SBP-BRiMS 2025
Y2 - 14 October 2025 through 17 October 2025
ER -