TY - GEN
T1 - On Causal and Anticausal LLM-based Data Synthesis
AU - Jiang, Bohan
AU - Ma, Pingchuan
AU - Shi, Zhuoyu
AU - Morstatter, Fred
AU - Raglin, Adrienne
AU - Liu, Huan
N1 - Publisher Copyright:
© 2026 Owner/Author.
PY - 2026/2/21
Y1 - 2026/2/21
N2 - While Large Language Models (LLMs) have been increasingly used to generate synthetic data for various downstream tasks, researchers overlook the causal direction in the data synthesis process. A natural causal direction should contain two steps: diverse raw data are generated first, and subsequently annotated for downstream tasks. However, most LLM-based methods adopt an anticausal direction: embedding label information in the prompt to force LLMs to generate targeted data. This reversal raises a critical question: How does the direction of data synthesis impact the quality and utility of the synthetic data? In this work, we empirically study the impact of causal and anticausal data synthesis. To do so, we first design simple yet effective prompting strategies to control the causal direction of LLM-based data synthesis. Using GPT-5 as the data generator, we construct synthetic datasets for three distinct machine learning tasks. We then fine-tune BERT-base and LLaMA-3.2-1B models on these datasets and evaluate them against human-curated benchmarks. Our experiments reveal consistent patterns: (1) models trained on anticausal synthetic data suffer larger performance drops across all tasks and model families - - Accuracy declines range from 13.7%-59.1% for BERT and 4.9%-54.3% for LLaMA, and (2) distributional analysis shows that anticausal synthetic datasets deviate further from human data. Our findings provide practical guidance on how to generate better synthetic data and make good use of it.
AB - While Large Language Models (LLMs) have been increasingly used to generate synthetic data for various downstream tasks, researchers overlook the causal direction in the data synthesis process. A natural causal direction should contain two steps: diverse raw data are generated first, and subsequently annotated for downstream tasks. However, most LLM-based methods adopt an anticausal direction: embedding label information in the prompt to force LLMs to generate targeted data. This reversal raises a critical question: How does the direction of data synthesis impact the quality and utility of the synthetic data? In this work, we empirically study the impact of causal and anticausal data synthesis. To do so, we first design simple yet effective prompting strategies to control the causal direction of LLM-based data synthesis. Using GPT-5 as the data generator, we construct synthetic datasets for three distinct machine learning tasks. We then fine-tune BERT-base and LLaMA-3.2-1B models on these datasets and evaluate them against human-curated benchmarks. Our experiments reveal consistent patterns: (1) models trained on anticausal synthetic data suffer larger performance drops across all tasks and model families - - Accuracy declines range from 13.7%-59.1% for BERT and 4.9%-54.3% for LLaMA, and (2) distributional analysis shows that anticausal synthetic datasets deviate further from human data. Our findings provide practical guidance on how to generate better synthetic data and make good use of it.
KW - causal and anticausal
KW - data synthesis
KW - large language models
UR - https://www.scopus.com/pages/publications/105033151506
UR - https://www.scopus.com/pages/publications/105033151506#tab=citedBy
U2 - 10.1145/3773966.3779380
DO - 10.1145/3773966.3779380
M3 - Conference contribution
AN - SCOPUS:105033151506
T3 - WSDM 2026 - Proceedings of the 19th ACM International Conference on Web Search and Data Mining
SP - 1160
EP - 1164
BT - WSDM 2026 - Proceedings of the 19th ACM International Conference on Web Search and Data Mining
PB - Association for Computing Machinery, Inc
T2 - 19th ACM International Conference on Web Search and Data Mining, WSDM 2026
Y2 - 22 February 2026 through 26 February 2026
ER -