INTEGRITY Cloudflare Docs

PII detection

AI Security for Apps (formerly Firewall for AI) can detect personally identifiable information (PII) in incoming LLM prompts. There are two approaches to PII detection, and you can use them together for layered protection:

AI-based PII detection

When AI Security for Apps is enabled and a request arrives at a cf-llm labeled endpoint, it scans the prompt for PII and populates two fields:

The detection is powered by an AI-based Named Entity Recognition (NER) model. Refer to the cf.llm.prompt.pii_categories field reference for the full list of recognized categories.

Supported PII categories

Category Description
BANK_ACCOUNT Bank account number
CREDIT_CARD Credit card number
DATE_TIME Date or time expression
DRIVER_LICENSE Driver license number
EMAIL_ADDRESS Email address
IP_ADDRESS IPv4 address
LOCATION Physical location or address
PASSPORT Passport number
PERSON Full or partial name of an individual
PHONE_NUMBER Phone number
TAX_ID Tax identification number
US_SSN US Social Security Number
URL URL

Be specific to reduce false positives

The cf.llm.prompt.pii_detected field returns true when any PII category is detected — including broad categories like PERSON, DATE_TIME, and LOCATION that frequently appear in normal conversation. Blocking based on this field alone will produce a high false-positive rate for most applications.

Instead, build rules against cf.llm.prompt.pii_categories and list only the categories that matter for your use case. For example, a customer support chatbot may need to block credit card numbers and SSNs but can safely ignore person names and dates. Start with the narrowest set of categories, monitor matches in Security Analytics, and expand only as needed.

Example rules — AI-based detection

Block any request containing PII

Block only specific PII categories

Log email addresses but block credit cards and SSNs

Create two custom rules:

  1. A rule with action Block and the following expression:
    (any(cf.llm.prompt.pii_categories[*] in {"CREDIT_CARD" "US_SSN"}))

  2. A rule with action Log and the following expression:
    (any(cf.llm.prompt.pii_categories[*] in {"EMAIL_ADDRESS"}))

Exact PII detection (regex)

If you need to detect custom PII formats specific to your organization — such as internal employee IDs, patient record numbers, or proprietary account identifiers — you can create a WAF custom rule using a regex match on the raw body (http.request.body.raw field).

This approach complements AI-based detection by matching predefined patterns, including organization-specific identifiers.

Example: Detect employee IDs

In the following example, an organization uses employee IDs in the format EMP- followed by exactly six digits (for example, EMP-482910).

Create a custom rule with the following configuration:

Scope to a specific endpoint

To limit this rule to only your LLM endpoint, combine it with a path condition:

Field Operator Value Logic
URI Path equals /api/chat And
Raw request body matches regex EMP-[0-9]{6}

Expression when using the editor:
(http.request.uri.path eq "/api/chat" and http.request.body.raw matches "EMP-[0-9]{6}")

More regex examples

Custom PII type Example format Regex pattern
Employee ID EMP-482910 EMP-[0-9]{6}
Patient record number PAT/2024/00391 PAT/[0-9]{4}/[0-9]{5}
Internal account ID ACCT-XX-99999 ACCT-[A-Z]{2}-[0-9]{5}
Custom API key prefix sk_live_abc123... sk_live_[a-zA-Z0-9]{20,}

Considerations for regex rules

Combine both approaches

You can use AI-based and exact detection together for layered protection:

(cf.llm.prompt.pii_detected or http.request.body.raw matches "EMP-[0-9]{6}")

This rule blocks requests where either the AI model detects any built-in PII category or the regex matches your custom identifier format.