Cisco Secure Access Help

PDF

Cisco Secure Access Help

Built-in Data Identifiers

Want to summarize with AI?

Log in

Describes Built-in Data Identifiers in Cisco Secure Access. Secure Access provides a wide selection of built-in data identifiers you can choose to incorporate into DLP policy to classify a wide range of sensitive data, both structured and unstructured.


Secure Access provides a wide selection of built-in data identifiers you can choose to incorporate into DLP policy to classify a wide range of sensitive data, both structured and unstructured.

The built-in identifier can classify specific data values based on pattern matching, bloom filters, and dictionaries incorporating proximity terms.

The ML (machine learning) built-in data identifiers are based on a LLM (Large Language Model) that was trained to classify unstructured sensitive data based on its true context. The ML data identifiers work for DLP-supported file types (see Supported File and Form Types); they do not work for form data. The ML built-in data identifiers are:

  • Bank Statement

  • Consulting Agreement

  • CV/Resume

  • Employment Agreement

  • IRS Forms

  • Medical Power of Attorney

  • Mergers And Acquisitions

  • NDA

  • Partnership Agreement

  • Source Code (ML)

  • Stock

  • US Patents

Built-in identifiers are not directly incorporated into DLP rules; you must first select and incorporate them into Data Classifications which you then apply to DLP rules. (See Manage Data Classifications ).

The built-in data identifiers are available as an Excel table here. The table is updated frequently, so be sure to download the most recent version.


Tolerances

Some built-in data identifiers have three versions, each with a different tolerance level associated with it. Tolerance levels determine how many instances of an identifier must appear within a document for that document to be considered a match. There are three tolerance levels:

  • Lenient tolerance requires only a single instance of an identifier to appear within a document for that document to be considered a match.
  • Moderate tolerance requires more occurrences of an identifier than Lenient tolerance, but fewer than Strict tolerance. (The exact number varies depending upon the identifier.)
  • Strict tolerance requires the most occurrences of an identifier to appear within a document for that document to be considered a match. (The exact number varies depending upon the identifier.)

For identifiers that support tolerance levels, you can choose the tolerance level that best matches your level of concern for catching every occurrence of those identifiers.


Copy and Customize a Data Identifier

You can copy a built-in data identifier and change the threshold and proximity on the copy to create your own customized data identifier. For a list of built-in data identifiers, see Built-in Data Identifiers.

Before you begin

Full Admin user role. For more information, see Manage Accounts.

Procedure

  1. Navigate to Secure > Settings > Data Classification.

  2. Expand a data classification, then expand a included data identifier within that classification and click COPY & CUSTOMIZE.

    Included Data Identifiers page displaying options for Copy and Customize
  3. Give the data identifier a meaningful name and description.

    Copy and Customize Data Identifier page displaying options for data identifier name and description
  4. (Optional) Select Threshold or Unique Threshold and then specify the Severity Criteria. The Threshold value represents the total number of occurrences of this identifier that must be detected in a document for Umbrella to generate an incident. The Unique Threshold value represents the number of unique occurrences of this identifier that must be detected in a document for Umbrella to generate an incident. A threshold of 10, for example, only generates an incident if 10 instances of the identifier are found in the file. The default threshold is 1.

    Copy and Customize Data Identifier page displaying an option for threshold

    To specify the Severity criteria:

    1. Select the severity name from the dropdown list. You can choose a predefined name or select Custom and then provide a name.
    2. Select the operator to define the value. You can select Range to define a range of the threshold value.
    3. Enter the threshold value or the range based on the selected operator.
      Copy and Customize Data Identifier page displaying options for threshold and unique threshold
  5. To add proximity keywords to match against the pattern, enter a proximity keyword in the proximity field and then click ADD. Repeat this step to add multiple proximity keywords. Secure Access will not generate an incident unless at least one of the occurrences of a matching term or pattern appears within 20 terms of a proximity keyword.

    Copy and Customize Data Identifier page displaying an option for proximity
  6. Click SAVE.

    The custom identifier appears under Custom Identifiers when creating or editing a classification. (See Create a Data Classification or Copy and Customize a Built-In Data Classification.)

What to do next

Note
Pattern

The pattern field is the built-in regular expression designed for this identifier and cannot be edited.


Create a Custom Identifier

A custom identifier can be used as part of a data classification to define terms and pattern expressions a DLP rule must match to generate an incident. You can optionally further restrict matching criteria using additional qualifiers:

Specify a threshold, to indicate the minimum number of occurrences of a term or pattern that must appear within a document to qualify as a match.

Specify proximity keywords, at least one of which must appear within 20 terms of a term or pattern to qualify as a match.

Custom identifiers can also be used to match text extracted from images when OCR scanning is configured in Global Settings.

Before you begin

Full Admin user role. For more information, see Manage Accounts.

Procedure

  1. Navigate to Secure > Settings > Data Classficiation.

  2. Click Add Custom Identifier.

  3. Give the custom identifier a meaningful name and description.

    Add Custom Identifier page displaying options for identifier name and description
  4. (Optional) Select Threshold or Unique Threshold and then specify the Severity Criteria. The Threshold value represents the total number of occurrences of this identifier that must be detected in a document for Secure Access to generate an incident. The Unique Threshold value represents the number of unique occurrences of this identifier that must be detected in a document for Secure Access to generate an incident. A threshold of 3, for example, only generates an incident if 3 instances of the identifier are found in the file. The default threshold is 1.

    Add Custom Identifier page displaying an option for threshold

    To specify the Severity criteria:

    1. Select the severity name from the dropdown list. You can choose a predefined name or select Custom and then provide a name.
    2. Select the operator to define the value. You can select Range to define a range of the threshold value.
    3. Enter the threshold value or the range based on the selected operator.
      Add Custom Identifier page displaying options for threshold and unique threshold
  5. (Optional) Specify a proximity keyword and then click ADD. You can repeat this step to add up to 10 proximity keywords for your custom identifier. If you specify proximity keywords, Secure Access will not generate an incident unless at least one of the occurrences of a matching term or pattern appears within 20 terms of a proximity keyword.

    Add Custom Identifier page displaying an option for proximity
  6. Add terms and patterns to your custom identifier.

    Note

    You can not upload or enter terms and patterns as a CSV. You must enter each item individually.

    1. For Terms, add up to 100 terms and click ADD for each term.
      Page displaying options for entry type, term, and add
    2. For Pattern, enter a regular expression 3-1,000 characters long. You can add up to 10 patterns to a custom identifier. For more information, see Custom Regular Expression Patterns.

      After adding a regular expression, you can enter sample text in the Test box, and click TEST to verify the pattern matches.

      Click ADD when you are satisfied that you have the pattern you want.

      Page displaying options for pattern and test

      Enter terms or patterns in the language you want DLP to detect, including text that may be extracted from images by OCR scanning.

  7. Click SAVE.

    Add New Data Classification page displaying an option for new banking terms

What to do next

If you specify both a threshold and proximity keywords, a document is considered a match only if it contains a custom identifier term or pattern that appears at least the number of times indicated by threshold, and at least one of those occurrences appears within 20 terms of at least one of the proximity keywords.

Note
Secure Access blocks documents containing a custom identifier only if the identifier is surrounded by a word boundary. A document containing a custom identifier with an alphanumeric character (a-z, A-Z, 0-9) adjacent to it will not be blocked. For example, for the custom identifier c.t matches cat and cot, but not housecat or citation.

If you are validating multilingual OCR scanning, you can create separate custom identifiers for each language and then add them to separate or combined data classifications.

Once you create a custom identifier, you can add it to a data classification. To create a data classification, see Create a Data Classification or Copy and Customize a Built-In Data Classification.


Custom Regular Expression Patterns

Note
You must create your own regex. The creation of custom regex is outside the scope of Secure Access Support.

Custom regular expression patterns for DLP data classifications support basic Java syntax with some limitations.

Limitations

General

  • A regex can have a maximum of 150 characters.
  • A custom dictionary can have a maximum of 10 regex.
  • A custom dictionary can have a maximum of 100 entries.
  • The minimum length of a regex pattern is 3.
  • The maximum number of matches for a regex is 1000.

Regex Syntax

  • Anchor Flags
    • ^ can not be used as an input start position.
    • $ can not be used as an input end position.
  • Back References
    • \n cannot be used to reference a previous capture group n.

The following must be used for these character definitions:

  • Whitespace—'\u0020'
  • Dash—'\u002D'
  • Single quote—(char)0x0027
  • Double quote—'\u0022'

Regex Breadth

  • Wildcards
    • .* and .+ are not accepted.
    • ()* and ()+ are not accepted.
  • | alternations have a maximum of 20 per pattern.

Word Boundary

The regular expression is wrapped in a word boundary to ensure it matches the entire word. Words are comprised of alphanumeric character (a-z, A-Z, 0-9) and the boundary cannot contain a word character on each side of the boundary.

Example:When the regular expression is cat, the sentence "The cat chased the mouse" will match, but "The catchased the mouse" will not match.


Individual Data Identifiers

These identifiers only appear in combination with other identifiers.

Drug Name

Drug Name is a content identifier that identifies unique drug names as referenced in the "Orange Book" (Approved Drug Products with Therapeutic Equivalence Evaluations) published by the US Food and Drug Administration (FDA). It is available from the FDA.

Health Condition

Health Condition is a content identifier that detects health conditions based on a comprehensive dictionary. The dictionary of health conditions is compiled from https://www.cdc.gov/nndss/index.html.

ICD-10 Code

The International Statistical Classification of Diseases and Related Health Problems (ICD) is a classification list compiled by the World Health Organization (WHO) for the purpose of coding diseases, signs and symptoms, abnormal findings, complaints, social circumstances, and external causes of injury or diseases. This policy identifies content in the 11th revision of ICD.

US Person Name

Person name is a content identifier used in numerous policies. It is backed by US Census information.