Cryptography-based privacy-preserving large language models: a lifecycle survey of frameworks, methods, and future directions
Cryptography-based privacy-preserving large language models: a lifecycle survey of frameworks, methods, and future directions
Artificial Intelligence Review 2026Artificial Intelligence Review 2026
基于 Transformer 的大语言模型正逐渐成为重要的技术基础设施,但其快速部署也使隐私泄露风险从局部问题扩展为贯穿模型全生命周期的系统性威胁,并限制了大语言模型在敏感及强监管行业中的规模化应用。
全同态加密、安全多方计算等密码学技术具有可证明的安全保障,已经应用到大语言模型的数据选择、微调和推理等关键阶段。然而,相关研究仍较为分散,缺少覆盖密码学隐私保护大语言模型全生命周期的综合梳理。
该综述系统回顾并分类现有框架与方法,为协调不同优化策略、设计更高效的隐私保护大语言模型算法提供结构化参考;同时结合现有研究的局限,归纳值得进一步探索的研究方向。
Transformer-based large language models are becoming critical technological infrastructure, but their rapid deployment has turned privacy leakage from a localized concern into a systemic risk spanning the entire model lifecycle. These risks restrict adoption in sensitive and heavily regulated domains.
Cryptographic techniques such as fully homomorphic encryption and secure multi-party computation offer provable security guarantees and are increasingly used in key stages including data selection, fine-tuning, and inference. Nevertheless, work on cryptography-based privacy-preserving LLMs remains fragmented.
This survey systematically reviews and classifies existing frameworks and methods, providing a structured basis for coordinating optimization strategies and designing efficient privacy-preserving LLM algorithms. It also identifies limitations in current research and outlines promising directions for future exploration.