What Data Can't Go Into an AI Tool
Understanding what data cannot go into AI tools is critical for sales teams to maintain compliance and security, preventing breaches and legal issues.
Avoid inputting sensitive data into AI tools
Sales teams must not enter PII, PCI, PHI, intellectual property, or confidential business info into AI tools without strong security and compliance assurances.
read: customer-data-compliance-in-ai-sales-tools/Assume all AI input data could become public
Treat any data entered into AI tools as potentially public unless the vendor guarantees data isolation, encryption, and no use for training.
read: ai-vendor-security-review-questions/Key data categories to keep out of AI tools
PII, PCI, PHI, intellectual property, confidential business info, and sensitive employee data should generally be excluded from AI inputs to avoid risks.
Some data can be used if anonymized or consented
Anonymized or pseudonymized data and publicly available data may be used with caution, and sensitive data only with explicit consent and secure handling.
Implement policies and training to manage AI data use
Create clear AI usage policies, conduct vendor security reviews, use data loss prevention tools, and train sales teams on data risks and compliance.
read: ai-usage-policy-template-sales/Unapproved AI use risks data exposure
Shadow AI occurs when employees use unapproved AI tools, risking exposure of confidential data; approved tools and education help prevent this.
Want this mapped to your stack?
30 minutes. We diagnose where your sales stack leaks and where AI actually fits. No vendor pitch.
Book a discovery callSales teams are increasingly adopting AI tools to enhance productivity, but a critical question often goes unanswered: what data cannot go into AI tools? The short answer is any data that poses a significant risk if exposed, misused, or if its use violates privacy regulations or company policy. This includes Personally Identifiable Information (PII), Payment Card Industry (PCI) data, protected health information (PHI), intellectual property, and confidential business information.
Using AI tools without clear data boundaries can lead to severe consequences, from regulatory fines to data breaches and loss of competitive advantage. It is not enough to simply adopt an AI tool; understanding and enforcing strict data governance is paramount.
The Core Principle: Assume Public Until Proven Private
When evaluating any AI tool, especially those that interact with large language models (LLMs), assume that any data you input could become public or be used to train the model for others. This conservative stance is vital unless the vendor explicitly guarantees data isolation, encryption, and non-use for training. Even then, verify these claims through a thorough security review, as outlined in our guide on AI vendor security review questions.
This principle applies to both third-party tools and internal AI solutions. While internal tools might offer more control, they still require robust data handling protocols.
Categories of Data to Restrict from AI Tools
Not all data is created equal. Here are the primary categories of information that sales teams should generally restrict from AI tools, along with the reasons why.
1. Personally Identifiable Information (PII)
PII includes any data that can be used to identify an individual. This is perhaps the most common type of sensitive data sales teams handle.
- Examples: Customer names, email addresses, phone numbers, physical addresses, social security numbers, employee IDs, IP addresses, unique device identifiers.
- Why restrict it:
- Privacy Regulations: Laws like GDPR, CCPA, and others impose strict rules on how PII is collected, processed, and stored. Unauthorized disclosure can lead to massive fines.
- Reputational Damage: Data breaches involving PII erode customer trust and can severely damage a company’s reputation.
- Identity Theft Risk: Exposure of PII can facilitate identity theft and other malicious activities.
Sales teams often use PII for outreach and CRM management. When integrating AI tools, ensure they are designed to handle PII securely, with appropriate data anonymization or pseudonymization features, or that the vendor’s data processing agreements explicitly cover PII protection. For more on this, see our article on customer data compliance in AI sales tools.
2. Payment Card Industry (PCI) Data
PCI data refers specifically to credit card information.
- Examples: Credit card numbers, expiration dates, CVV codes, cardholder names.
- Why restrict it:
- PCI DSS Compliance: The Payment Card Industry Data Security Standard (PCI DSS) is a global standard for organizations that handle branded credit cards. Non-compliance can result in severe penalties, fines, and loss of ability to process card payments.
- Financial Fraud: Exposure of PCI data directly leads to financial fraud.
Sales teams should never input PCI data into any AI tool, regardless of perceived security. Payment processing should always occur through certified, secure payment gateways that are PCI compliant.
3. Protected Health Information (PHI)
PHI is any health information about an individual that is created, received, stored, or transmitted by a HIPAA-covered entity.
- Examples: Medical records, health insurance information, patient names, dates of birth, treatment histories.
- Why restrict it:
- HIPAA Compliance: The Health Insurance Portability and Accountability Act (HIPAA) in the U.S. mandates strict privacy and security rules for PHI. Similar regulations exist globally.
- Ethical Concerns: Health data is extremely sensitive and its misuse can have profound personal consequences.
While less common for general sales teams, those in healthcare sales must be acutely aware of PHI restrictions. AI tools used in this sector require specific HIPAA-compliant certifications and data handling protocols.
4. Intellectual Property (IP) and Confidential Business Information (CBI)
This category covers proprietary company data that provides a competitive edge.
- Examples: Unreleased product roadmaps, proprietary algorithms, unique sales methodologies, market research data, M&A plans, financial forecasts, trade secrets, confidential client lists (beyond basic contact info).
- Why restrict it:
- Competitive Advantage: Exposure of IP or CBI can undermine a company’s market position, allowing competitors to replicate strategies or products.
- Loss of Value: Trade secrets lose their value once they are no longer secret.
- Legal Ramifications: Breaches of confidentiality agreements can lead to lawsuits.
Sales teams might be tempted to use AI to summarize internal strategy documents or analyze proprietary sales playbooks. This is a high-risk activity. If an AI tool uses input data for training, your confidential information could inadvertently become part of the public model’s knowledge base or be exposed to other users.
If you wouldn’t email it to a competitor, don’t put it into an unverified AI tool.
5. Sensitive Employee Data
While often managed by HR, sales leaders might have access to certain employee data.
- Examples: Employee performance reviews, salary information, disciplinary records, personal health information of employees.
- Why restrict it:
- Privacy Laws: Many jurisdictions have laws protecting employee privacy.
- Internal Trust: Breaching employee confidentiality can severely damage morale and trust within the organization.
The Nuance: When Data Can Be Used (With Caution)
Not all data is off-limits. The key is understanding the AI tool’s architecture, the vendor’s data policies, and your organization’s risk tolerance.
Anonymized or Pseudonymized Data
If sensitive data can be effectively anonymized (stripped of all identifying information) or pseudonymized (identifying information replaced with a reversible code), it may be suitable for AI analysis.
- Use Case: Training an AI model on sales call transcripts to identify common objections, where speaker identities are removed.
- Caveat: True anonymization is complex. Ensure the process is robust enough to prevent re-identification, especially when combining multiple datasets.
Publicly Available Data
Data that is already public can generally be used freely.
- Use Case: Analyzing public company reports, industry news, or social media sentiment to inform sales strategies.
- Caveat: Even public data can have usage restrictions (e.g., copyright). Always respect terms of service.
Data with Explicit Consent and Secure Handling
In specific cases, with explicit, informed consent from individuals and a highly secure, compliant AI tool, certain sensitive data might be used.
- Use Case: A customer agrees to have their anonymized feedback used to train an AI model that improves product features.
- Caveat: This requires robust legal frameworks, transparent communication, and verifiable security measures from the AI vendor.
Practical Steps for Sales Teams
To manage data boundaries effectively, sales teams need clear policies and processes.
1. Develop a Comprehensive AI Usage Policy
This policy should clearly define what data can and cannot be used with AI tools, which tools are approved, and the consequences of non-compliance. Our guide on an AI usage policy template for sales offers a starting point.
| Data Type | General Guideline | Risk Level |
|---|---|---|
| PII (Customer Names, Emails) | Restrict unless anonymized or with explicit secure vendor DPA | High |
| PCI Data (Credit Cards) | Absolutely Prohibited | Critical |
| PHI (Health Records) | Absolutely Prohibited (unless HIPAA-compliant tool) | Critical |
| Internal IP (Roadmaps) | Restrict unless isolated and non-training guaranteed | High |
| Public Company Data | Generally Permitted | Low |
| Anonymized Sales Data | Permitted with robust anonymization | Medium |
2. Conduct Vendor Security Reviews
Before adopting any new AI tool, perform a thorough security and compliance review. Ask vendors specific questions about their data handling, encryption, data residency, and whether input data is used for model training.
3. Implement Data Loss Prevention (DLP) Tools
DLP solutions can help prevent sensitive data from being inadvertently uploaded to unapproved AI platforms. These tools can scan outgoing data for PII, PCI, and other confidential information.
4. Train Your Team
Regular training is crucial. Sales reps need to understand why certain data is restricted and the potential consequences of non-compliance. Make the policy accessible and easy to understand.
5. Prioritize CRM Data Hygiene
Poor data quality in your CRM can exacerbate AI data risks. If your CRM contains inaccurate or improperly classified sensitive data, it increases the chance of it being fed into an AI tool where it shouldn’t be. Focusing on CRM data hygiene before AI is a foundational step.
The Shadow AI Problem
One of the biggest challenges is “shadow AI,” where employees use unapproved AI tools without company oversight. This often happens because employees are trying to be efficient but are unaware of the data risks.
- Scenario: A sales rep copies a confidential client proposal into a public AI chatbot to summarize it or improve its language.
- Risk: The proposal’s contents, including client-specific details or proprietary strategies, could be absorbed by the public model and potentially exposed.
Addressing shadow AI requires a combination of clear policies, approved tools, and ongoing education. It’s about empowering your team to use AI safely, not just restricting its use.
Shadow AI isn’t just a security risk; it’s a symptom of unmet productivity needs. Provide approved, secure tools, or your team will find their own.
Conclusion
The promise of AI in sales is immense, but its responsible adoption hinges on strict data governance. Understanding what data cannot go into AI tools is not just a compliance exercise; it’s a fundamental aspect of protecting your company’s assets, reputation, and customer trust. By implementing clear policies, vetting vendors, and continuously educating your sales team, you can harness the power of AI while mitigating its inherent risks.
FAQ
What are the primary risks of putting sensitive data into AI tools?
The primary risks include data breaches, non-compliance with regulations like GDPR or CCPA, loss of intellectual property, and reputational damage. AI models can inadvertently expose or misuse sensitive information if not properly controlled.
How does PII differ from PCI data in the context of AI tools?
Personally Identifiable Information (PII) refers to data that can identify an individual, such as names, emails, or addresses. Payment Card Industry (PCI) data specifically relates to credit card information. Both are highly sensitive, but PCI has stricter handling requirements due to financial fraud risks.
Can internal strategy documents be shared with AI tools?
Generally, no. Internal strategy documents, especially those containing unreleased product roadmaps, M&A plans, or competitive intelligence, should not be shared with public or unsecure AI tools. This risks intellectual property leakage and competitive disadvantage.
What is the role of data anonymization in using AI tools safely?
Data anonymization transforms sensitive data so that individuals cannot be identified, allowing it to be used for training or analysis by AI tools without privacy risks. It's a key technique for leveraging AI while protecting confidentiality.
Why is it important to have an AI usage policy for sales teams?
An AI usage policy provides clear guidelines on what data can and cannot be used with AI tools, defines approved tools, and outlines compliance procedures. This prevents shadow AI, reduces risk, and ensures consistent, secure AI adoption across the sales organization.
Want a stack audit instead of another vendor pitch? Book a discovery call.
Book a discovery call

