Skip to main navigation Skip to search Skip to main content

Can Typos Cause Harm? The Impact of Imperfect Input on LLM Safety

  • Saurabh Zinjad
  • , Amrita Bhattacharjee
  • , Alimohammad Beigi
  • , Huan Liu

Research output: Chapter in Book/Report/Conference proceedingConference contribution

Abstract

Large Language Models (LLMs) are increasingly being used in sensitive domains such as healthcare and education, where safety is critical. While much LLM safety research focuses on deliberate attacks (e.g., jailbreaks, prompt injections), the impact of benign but imperfect user input, such as typos or paraphrasing remains underexplored. In this study, we investigate how such semantic-preserving perturbations affect the safety behavior of aligned LLMs. We find that these small perturbations can increase instability, e.g., causing unsafe responses to flip into safe refusals. This bidirectional instability reveals that current safety alignment mechanisms are fragile and context-dependent. Using a wide set of realistic perturbations in harmful queries, we evaluated several open-source models. Our results highlight the need for more context-aware and model-sensitive evaluation frameworks and training methods that ensure robust behaviors in the face of natural and noisy input. Github: https://github.com/Ztrimus/llm-sensitivity.

Original languageEnglish (US)
Title of host publicationSocial, Cultural, and Behavioral Modeling - 18th International Conference, SBP-BRiMS 2025, Proceedings
EditorsRobert Thomson, Aryn A Pyke, Scott Renshaw, Patrick Park, Samer Al-khateeb, Annetta Burger
PublisherSpringer Science and Business Media Deutschland GmbH
Pages233-243
Number of pages11
ISBN (Print)9783032077141
DOIs
StatePublished - 2026
Event18th International Conference on Social Computing, Behavioral-Cultural Modeling and Prediction and Behavior Representation in Modeling and Simulation, SBP-BRiMS 2025 - Pittsburgh, United States
Duration: Oct 14 2025Oct 17 2025

Publication series

NameLecture Notes in Computer Science
Volume16127 LNCS
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference18th International Conference on Social Computing, Behavioral-Cultural Modeling and Prediction and Behavior Representation in Modeling and Simulation, SBP-BRiMS 2025
Country/TerritoryUnited States
CityPittsburgh
Period10/14/2510/17/25

Keywords

  • Input Sensitivity
  • Prompt Perturbations
  • Robustness Evaluation
  • Safety Alignment
  • Safety Flip Rates

ASJC Scopus subject areas

  • Theoretical Computer Science
  • General Computer Science

Fingerprint

Dive into the research topics of 'Can Typos Cause Harm? The Impact of Imperfect Input on LLM Safety'. Together they form a unique fingerprint.

Cite this