SaaS AI apps के लिए prompt injection से बचाव
Trusted boundaries, least privilege, safe retrieval, output checks और adversarial tests से SaaS AI में direct तथा indirect prompt injection का जोखिम घटाएं।
इस मार्गदर्शिका में
SaaS AI feature को prompt injection से कैसे बचाएं?
Prompt injection तब होता है जब user input या बाहरी सामग्री language model के व्यवहार को बदलने की कोशिश करती है। Direct attack user के prompt में आता है; indirect attack webpage, document, email या RAG result में छिपा हो सकता है। कोई prompt wording या filter इसे पूरी तरह रोकने की गारंटी नहीं देता। Feature इस तरह बनाएं कि injection सफल होने पर भी model access न दे सके, secrets न बताए और प्रभावशाली actions को authorize न कर सके।
Authorization और secrets को model prompt से बाहर रखें
Model तक data पहुंचने से पहले application code में tenant membership, object access और business permissions लागू करें। API keys, system credentials, छिपा cross-tenant data या unrestricted internal instructions prompt में न डालें। Context में भेजी हर चीज़ को संभावित रूप से दोहराया या बदला जा सकने वाला मानें; task के लिए जरूरी न्यूनतम authorized जानकारी ही भेजें।
User और retrieved content को untrusted data मानें
Trusted task instructions को quoted user text, retrieved passages और tool results से संरचनात्मक रूप से अलग रखें। Prompt में boundaries बताएं और source labels रखें, लेकिन इन्हें सहायक संकेत समझें, security control नहीं। Document का instruction server policy, identity check या तय tool permission list को बदल न सके।
सफल injection से होने वाला नुकसान सीमित करें
Model को वर्तमान feature के लिए जरूरी tools ही दें, जहां संभव हो read-only access रखें, हर operation को मौजूदा user और tenant तक सीमित करें, और संवेदनशील या irreversible action से पहले मानव की मंजूरी लें। Rate, value और destination limits service layer में लागू हों। Refusal prompt इन enforceable सीमाओं का विकल्प नहीं है।
| Feature और untrusted input | Model को दिखने वाला data | Tools और server checks | Approval या सीमा | Adversarial test owner |
|---|---|---|---|---|
| Uploaded files का सार | ||||
| Tenant knowledge base से जवाब | ||||
| बाहरी message का draft |
Direct और indirect prompt injection के लिए कौन से controls मददगार हैं?
Input validate करें और task को सीमित रखें
Model processing से पहले size, file type, encoding और request rate की सीमा रखें। Narrow task instructions, अनुमत output formats और feature से बाहर requests के लिए स्पष्ट refusal behavior तय करें। Input filters ज्ञात patterns पकड़ सकते हैं, लेकिन हमलावर paraphrase, translation, encoding या सामान्य content में निर्देश छिपा सकते हैं; keyword blocking को पूरा बचाव न मानें।
Web pages और documents के retrieval में exposure घटाएं
केवल जरूरी sources fetch करें, जहां उचित हो active markup हटाएं, provenance रखें और trusted developer instructions को source text से न मिलाएं। URL fetching SSRF-safe controls के पीछे रखें; retrieved page को आगे की request का tenant, query scope या credentials चुनने न दें। किसी document को पढ़ने की अनुमति होने से उसमें embedded commands भरोसेमंद नहीं हो जाते।
संवेदनशील सीमाओं पर model output को प्रस्ताव मानें
Email भेजने, account settings बदलने, records export करने या workflow update करने से पहले प्रस्तावित action को मौजूदा server-side policy और authenticated user की permissions के विरुद्ध validate करें। जहां confirmation चाहिए, user को intended recipient, records और परिणाम दिखाएं। Approval और नतीजा दर्ज करें, पर अनावश्यक prompt content न रखें।
SaaS team prompt injection बचाव का मूल्यांकन कैसे करे?
हर feature के लिए threat-based attack set बनाएं
Policy अनदेखा करने वाले direct requests, PDFs और webpages के indirect instructions, multilingual और obfuscated text, tool-result manipulation, लंबा context भटकाना और cross-tenant data निकालने की कोशिश जांचें। Benign lookalikes भी शामिल करें ताकि feature बेवजह unusable न हो। Prompt, tools, model, retrieval या provider बदलने पर tests दोबारा चलाएं।
Model के शब्द नहीं, security outcome जांचें
Test तभी pass हो जब unauthorized data unavailable रहे, निषिद्ध tool call न चले और approval bypass न हो। केवल assistant ने refusal कहा या नहीं, इतना न देखें; server authorization और side effects जांचें। Retries, streaming, fallback models और parallel tool calls भी test करें क्योंकि हर रास्ते में अलग जोखिम हो सकता है।
Abuse पर निगरानी और सुरक्षित failure path रखें
Feature version, tenant pseudonym, policy decision, tool name और outcome को तय retention के साथ log करें। Boundary probes, असामान्य retrieval breadth और denied tool calls पर alert दें। Safety check या provider उपलब्ध न हो तो privileged action रोक दें और बताएं कि task सुरक्षित रूप से पूरा नहीं हुआ।
SaaS prompt injection: अक्सर पूछे जाने वाले सवाल
क्या system prompt prompt injection को पूरी तरह रोक सकता है?
नहीं। Prompt instructions व्यवहार सुधार सकती हैं, पर hard security boundary नहीं बनातीं। Authorization, secrets, tool permissions और consequential action checks trusted application code में रखें।
क्या retrieval-augmented generation injection से सुरक्षित है?
नहीं। Retrieved document में malicious instructions हो सकते हैं। Retrieval अलग से authorize करें, source text को untrusted मानें, model access सीमित करें और indirect attacks जांचें।
क्या 'ignore previous instructions' जैसे phrases block करने चाहिए?
Pattern filter एक detection layer हो सकता है, लेकिन paraphrase और encoded content छूट सकते हैं और harmless text भी block हो सकता है। Filter attack न पकड़े तब भी feature सुरक्षित रहे, ऐसी design करें।
AI feature को human approval कब मांगना चाहिए?
जब action बाहरी हो, असर बड़ा हो, वापस लेना कठिन हो या access, पैसा, customer records अथवा कानूनी commitments बदलता हो। Reviewer को पूरी कार्रवाई दिखाएं और server से फिर जांचें कि उसे अनुमति है।
संबंधित व्यावहारिक मार्गदर्शिकाएँ
संबंधित मुद्दों की मार्गदर्शिकाएँ
स्रोत और प्रकाशन रिकॉर्ड
मसौदा 27 सितंबर 2026 को तैयार; इंजीनियरिंग, सुरक्षा और संपादकीय समीक्षा बाकी है · स्रोत जाँचे गए .
- LLM01:2025 Prompt Injection
- LLM08:2025 Vector and Embedding Weaknesses
- LLM06:2025 Excessive Agency
- OWASP Top 10 for LLM Applications 2025
- Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1)
- OWASP Cheat Sheet: authorization
- OWASP Cheat Sheet: logging
- OWASP Cheat Sheet: file upload
- OWASP Server-Side Request Forgery Prevention Cheat Sheet