Data Retention Policy: How to Write One That Holds
On January 13, 2026, the French data protection authority fined two telecom operators in the same group a combined 42 million euros. An attacker had reached data tied to more than 24 million subscriber contracts. One of the three violations had nothing to do with the attack itself: the companies had no way to sort the data of former subscribers and keep only what accounting rules required. They had a breach problem because they had a keeping problem. A data retention policy is the document that decides what you keep, for how long, why, and who deletes it. This guide shows how to write one that survives contact with real systems.
I write as a founder who has rebuilt processes in companies from private healthcare to distribution. Retention is the topic where I see the widest gap between paper and practice. Most companies have a policy. Very few have ever deleted anything because of it.
What a data retention policy is, and what it is not
A data retention policy is a set of rules that assigns a retention period, a legal or business reason, an owner and a disposal method to every category of record the company holds. It covers personal data and non-personal records alike: contracts, invoices, payroll files, emails, logs, customer accounts, candidate résumés, backups.
It has two parts that are often confused. The policy is the short document that states principles, roles and exceptions. The retention schedule is the table that lists each record category with its period and its trigger. The policy changes rarely. The schedule changes whenever the business or the law does.
A data retention policy is not a backup policy. Backups exist to restore systems after a failure. Retention exists to decide how long information should live at all. When the two are mixed, companies end up keeping ten years of backups "for compliance" and can no longer delete anything.
It is not a privacy notice either. The notice tells people how long you keep their data. The policy is the internal mechanism that makes the notice true. Regulators compare the two, and the distance between them is where fines begin.
Why retention became a board topic
Three forces pushed retention from the records room to the board agenda.
Breaches cost more when you keep more
IBM's 2026 Cost of a Data Breach study, based on 602 organizations breached between March 2025 and February 2026, puts the global average cost of a breach at 4.99 million dollars. According to the IBM announcement of the study, one in four malicious breaches was enabled by artificial intelligence, and those cost about 6 million dollars on average.
The same study found that only 37 percent of breached organizations encrypt sensitive data both at rest and in transit. Put the two facts together. Most companies hold large volumes of data with partial protection. Every record kept past its useful life adds to the damage of an incident and adds nothing to revenue.
Regulators now enforce storage limitation
European regulators issued roughly 1.2 billion euros in fines in 2025, and about 7.1 billion since 2018, according to the annual DLA Piper survey as reported by Bitdefender. Breach notifications rose 22 percent year on year and passed 400 per day on average for the first time.
Retention failures appear in these decisions again and again, usually next to a security failure. The CNIL decision against the two French operators is the clearest recent example: weak remote access controls opened the door, and years of former customer data made the room behind it much larger.
Artificial intelligence reads everything you forgot
Internal assistants and search tools index whatever they can reach. A shared drive full of old HR files, expired contracts and customer exports was a dormant risk when nobody opened it. Once an assistant can summarize it on request, it is an active one.
Companies that deploy enterprise search software discover this in the first week. The tool works, and it surfaces documents that should have been deleted years ago. Retention is now a precondition for using artificial intelligence on internal data, not a legal afterthought.
The legal floor: what the rules actually require
A retention schedule sits between two kinds of rules. Some set a minimum: keep this for at least a given period. Others set a maximum: do not keep this longer than you need. The policy has to satisfy both at once.
Rules that set a maximum
Under the GDPR, the storage limitation principle in Article 5 requires that personal data be kept in a form that permits identification for no longer than necessary for the purposes of the processing. The regulation gives no fixed periods. The company must decide them and be able to justify them.
The UK regulator is explicit about documentation. The ICO guidance on storage limitation states that you need a policy setting standard retention periods wherever possible, and that you should periodically review the data you hold and erase or anonymize it when you no longer need it. Small organizations with occasional low risk processing may not need a documented policy, but must still review and delete.
In the United States there is no single federal rule, but state privacy laws have moved in the same direction. California requires businesses to tell consumers how long each category of personal information will be kept, or the criteria used to decide, and prohibits keeping it longer than reasonably necessary for the disclosed purpose.
Rules that set a minimum
Minimum periods come from tax, employment, financial and sector rules. A few examples that apply to most American businesses:
| Record type | Minimum period | Source |
|---|---|---|
| Tax records supporting a return | 3 years in the standard case, 6 or 7 in specific situations | IRS |
| Employment tax records | At least 4 years after the tax is due or paid | IRS |
| Personnel and employment records | 1 year, or 1 year from involuntary termination | EEOC |
| Payroll records | 3 years | EEOC, under ADEA and FLSA |
| Audit and review workpapers of public company auditors | 7 years | SEC Regulation S-X |
| HIPAA required documentation, such as policies and procedures | 6 years | HIPAA Security Rule |
The IRS guidance on how long to keep records lists the cases: three years as the baseline, six years if unreported income exceeds 25 percent of gross income shown on the return, seven years for a claim involving worthless securities or a bad debt deduction, and indefinitely if no return was filed or the return was fraudulent.
The EEOC recordkeeping requirements add an important trigger. When a charge of discrimination has been filed, the relevant records must be kept until final disposition of the charge or any lawsuit. The period is no longer a number. It is an event.
The third rule: legal holds
When litigation is pending or reasonably anticipated, the duty to preserve overrides the schedule. Deleting under the policy after that point can be treated as spoliation. Every data retention policy needs a legal hold procedure that can suspend deletion for a defined set of records, and release it when the matter closes.
This is the clause that most templates include and most companies cannot execute. Suspending deletion requires knowing where the records are and which automated jobs touch them.
How to write a data retention policy for a business: eight steps
The order matters. Companies that start from step five, the periods, produce a table that looks complete and cannot be applied.
Step 1: Appoint an owner with authority
Retention crosses legal, IT, finance, HR and sales. It needs one accountable owner who can convene all of them. In mid sized companies this is usually the general counsel, the head of compliance or the CFO. IT is essential but should not own the policy, because the decisions are legal and commercial.
Step 2: Inventory records by category, not by system
List what the company holds in business terms: customer contracts, supplier invoices, candidate files, support tickets, call recordings, marketing lists, access logs. Thirty to sixty categories are enough for most companies. More than a hundred and nobody will maintain it.
Then map each category to the systems where it lives. One category usually lives in several places: the application, an export on a shared drive, an email attachment, a backup. The map is what makes deletion possible later.
Step 3: Identify the purpose and the legal basis
For each category, write one sentence on why the company holds it. If nobody can write the sentence, the category is a candidate for deletion, not for a retention period.
Step 4: Collect the legal minimums and maximums
For each category, list the rules that apply in each jurisdiction where the company operates. Record the source. A period without a source cannot be defended and will not be updated when the law changes.
Step 5: Set the period and the trigger
A period has two parts: a duration and a starting event. "Seven years" means nothing until you say seven years from what. Common triggers are the end of the contract, the end of employment, the end of the fiscal year, the last customer activity, or the closing of a ticket.
Triggers are where schedules fail in practice. If the system does not record the triggering event as a date, the period cannot be automated. Check this before approving the schedule.
Step 6: Define the disposal method
Deletion, anonymization and archiving with restricted access are three different outcomes. State which one applies to each category, and what evidence is kept. A disposal log that records what was deleted, when, under which rule and by whom is the proof that the policy is applied.
Step 7: Write the exceptions
Legal holds, regulatory investigations, ongoing disputes with a customer, and records of historical value. Each exception needs an approver and an expiry date. Exceptions without an expiry become the new default.
Step 8: Approve, publish and schedule the review
The policy should be approved at executive level, published where employees can find it, and reviewed at least once a year. The schedule should also be reviewed whenever a new system is introduced. Tying this review to vendor onboarding is the simplest control: no new vendor that stores company data goes live without a retention period and an exit clause.
The retention schedule: a working model
The schedule is the part people actually use. Keep it to one table with the same columns for every row.
| Record category | Owner | Trigger | Period | Source | Disposal | Systems |
|---|---|---|---|---|---|---|
| Customer contracts | Legal | Contract end | Set by limitation period in governing law | Contract law, tax rules | Delete, keep log | CRM, contract repository |
| Supplier invoices | Finance | Fiscal year end | Set by tax rules | Tax authority | Delete | ERP, document archive |
| Payroll records | HR | Payment date | At least 3 years in the US | EEOC, FLSA | Delete | Payroll system |
| Candidate files, not hired | HR | Decision date | At least 1 year in the US | EEOC | Delete | Applicant tracking system, email |
| Support tickets | Customer service | Ticket closed | Business decision | Internal | Anonymize | Help desk |
| Marketing contacts, inactive | Marketing | Last activity | Business decision, within privacy law | Privacy law | Delete | Marketing platform |
| Access logs | IT security | Log date | Business decision, within security rules | Security standard | Delete | Log platform |
Two notes on this model. The periods marked as business decisions are the hard ones, because no law gives the number. The company has to decide how long an inactive contact is still a prospect, or how long a closed ticket is still useful. Those decisions belong to the business owner, with legal advice, and they should be written down with the reasoning.
The second note concerns the last column. A schedule without the systems column is a statement of intent. With it, the schedule becomes a work order for IT.
Periods are per jurisdiction
A company with employees in the United States and customers in Europe holds the same record category under different rules. The practical approach is to set the period per jurisdiction where the difference is material, and to use the longest minimum only when it does not breach a maximum elsewhere. Applying the longest period everywhere is simple and often unlawful for personal data of European residents.
Where retention policies fail: seven patterns
The policy has no systems. It lists periods and no location. IT cannot act on it, so nothing is deleted.
The trigger does not exist as data. The policy says seven years after contract end, and the CRM has no contract end date. Someone has to add the field, and fill it for the past.
Backups are forgotten. Data deleted from production lives on in backups for years. The policy needs a rule: backups rotate on a fixed cycle, and restored data is subject to the schedule again.
Email is treated as out of scope. Email is the largest unstructured archive in most companies and holds copies of nearly every record category. A policy that ignores it covers half the data.
Exports and personal copies. Spreadsheets exported from the CRM, files on laptops, shared drives. Technical controls on exports matter more than policy text here.
Employee departures. Mailboxes and personal drives of former employees are kept "just in case" for years. The offboarding step in a hire to retire process should include a review date and a deletion date for both.
Nobody is measured on it. If no executive sees a number every quarter, the policy decays. What gets reported gets done.
The French decision shows several of these at once. The operators needed some data of former subscribers for accounting purposes. What they lacked was the ability to separate that data from everything else. The failure was architectural, not legal.
Deletion in practice: from policy to systems
Writing the policy is a legal exercise. Applying it is an engineering one. Four approaches exist, and most companies need a mix.
Native retention settings. Most business applications, email platforms and collaboration suites have built in retention rules. They are the cheapest way to enforce the schedule. Start here.
Scheduled jobs. For internal databases and older applications, deletion runs as a scheduled job written by IT. Each job needs an owner, a test, and an entry in the disposal log.
Anonymization. When the business needs the statistics but not the identity, anonymization keeps the value and removes the risk. It must be irreversible to count. Replacing names with an identifier that can be mapped back is pseudonymization, and the data is still personal.
Manual review. For unstructured files and edge cases, a periodic manual review is unavoidable. Keep it small by fixing the sources: fewer exports, fewer personal copies.
Test deletion before you trust it
A deletion job that has never run in production is a hypothesis. Run it first on a copy. Check what it removes, what depends on the removed data, and whether reports still work. In one hotel group I worked with, revenue grew from 9 million to 10 million after we rebuilt how customer data fed pricing and marketing. Part of that work was cleaning years of duplicate and outdated guest profiles. The analysis improved because the data was smaller and accurate.
Third parties hold your data too
Processors, cloud providers and software vendors hold copies of company data. The contract should state what happens at the end of the relationship: return or deletion, within how many days, with what certification. It should also state the vendor's own backup cycle. Without it, your schedule stops at the boundary of your network.
Data retention and artificial intelligence
Artificial intelligence adds three retention questions that older policies do not answer.
Prompts and outputs. Conversations with assistants are records. They can contain personal data, confidential information and business decisions. The policy should state how long they are kept and whether the vendor keeps them too.
Training and fine tuning data. If the company trains or adapts a model on its own data, deleting the source record does not remove what the model learned. The practical answer today is to control what goes in, and to document it.
Indexes and embeddings. Search tools build indexes of documents. When a document is deleted, the index must be updated. Ask the vendor how long it takes.
Unapproved tools make all three worse, because the data leaves the company without anyone deciding how long it stays outside. A program to manage shadow AI and a retention policy are two halves of the same control. Companies building an AI governance function should put retention on its first agenda.
Roles: who decides, who executes, who checks
A data retention policy fails when everyone is involved and nobody is accountable. Four roles are enough, and each needs a name next to it.
The policy owner approves the policy, arbitrates conflicts between functions and reports to the executive team. This person does not set every period. They make sure every period has been set by someone with the authority to set it.
Record owners are the heads of the functions that create the records: finance for invoices, HR for personnel files, sales for customer accounts. They decide the business periods, confirm the triggers and sign off on disposal. They are also the ones who will resist deletion, which is why their sign off has to be explicit.
System owners sit in IT. They translate the schedule into retention settings and scheduled jobs, maintain the disposal log and report on overdue volume. They execute decisions. They should not be asked to make them.
The reviewer is internal audit, compliance or an external advisor. Once a year they sample a few categories and check whether records past their period still exist. A sample of five categories is enough to show whether the program is real.
The conversation that always happens
At some point a record owner will say that the data might be useful one day. It is the most common objection and it deserves a serious answer. Ask three questions. When was this data last used for a decision? What would it cost to recreate it if needed? What would it cost if it were exposed? In most cases the first answer is "never" and the discussion ends there.
When the answer is that the data does feed analysis, anonymization is usually the right outcome. The business keeps the trend, and the company stops holding the identity.
The hard categories: email, chat, recordings and logs
Structured records in business applications are the easy part. Four categories cause most of the trouble.
Email. Mailboxes hold contracts, invoices, personal data and negotiations, mixed together. Classifying each message is unrealistic. The workable approach is a default retention period for mailboxes, with a defined way to move records that must be kept longer into the system where they belong. The contract goes to the contract repository. The mailbox copy expires.
Chat and collaboration tools. Messages are informal, and they are records. Decisions get made in them. Set a default period by channel type, shorter for direct messages and longer for project channels, and tell employees what it is.
Call and meeting recordings. Recording and transcription have become a default setting in many tools. Each recording contains the voices and statements of people who may not work for the company. Decide who can record, where recordings are stored and when they expire. The transcript needs the same period as the recording.
Logs. Security teams want logs for as long as possible, and logs contain personal data. Separate security logs needed for investigation from application logs kept out of habit. Give each a period and reduce the personal data in them where the purpose allows it.
In a medical center I worked with, capacity grew by 20 percent after we redesigned patient flows and the systems behind them. One finding was that staff spent time searching through years of duplicated patient documents stored in several places. Fixing where records lived, and for how long, removed both a privacy exposure and a daily waste of time.
What keeping data really costs
Storage is cheap, and that is the argument most often used against deletion. Storage is also the smallest line in the cost of keeping data.
| Cost line | When it appears | Who pays |
|---|---|---|
| Storage and backup | Every month | IT budget |
| Search and discovery in disputes | When litigation starts | Legal budget |
| Response to access and deletion requests | Every request | Operations |
| Breach impact | When an incident happens | The whole company |
| Migration of old data to new systems | Every system change | Project budget |
| Poor data quality in analysis | Every report | Every decision maker |
Discovery is the line that surprises executives. In a dispute, the company must search what it holds. The more it holds, the more it has to review, and legal review is paid by the hour. Data that should have been deleted under the policy still has to be produced if it exists.
Migration is the line that surprises IT. Every change of system raises the question of what to move. Without a schedule, the answer is everything, and the project carries fifteen years of records into a platform designed for the next five.
Self-assessment: does your policy hold?
Score one point for each statement that is fully true today. Partial answers count as zero.
| Number | Statement | Point |
|---|---|---|
| 1 | We have a written data retention policy approved in the last 24 months | |
| 2 | We have a retention schedule with a period and a trigger for each record category | |
| 3 | Every period has a documented legal or business source | |
| 4 | Each category is mapped to the systems where it lives | |
| 5 | At least one automated deletion rule is running in production | |
| 6 | We keep a disposal log | |
| 7 | We have a legal hold procedure, and we have tested it | |
| 8 | Backups rotate on a defined cycle that is consistent with the schedule | |
| 9 | Mailboxes and drives of former employees are deleted on a defined date | |
| 10 | Vendor contracts state what happens to our data at termination | |
| 11 | Our privacy notice and our schedule state the same periods | |
| 12 | An executive sees retention metrics at least twice a year | |
10 to 12 points. The policy is operating. Focus on automation coverage and on new systems, especially artificial intelligence tools.
6 to 9 points. You have the document and part of the mechanism. The gap is usually in statements 4, 5 and 8: systems, automation, backups. This is an IT project with a legal sponsor.
0 to 5 points. You have a policy on paper, or none. Start with the inventory and with one category where deletion is easy and the volume is large, such as inactive marketing contacts or candidate files.
If you score below six and cannot see where to start, send a consultation request describing your sector, your systems and the jurisdictions where you operate. The first conversation is about order of priority, not about tools.
Metrics that show whether the policy is alive
Five numbers are enough. Report them to the executive team every quarter.
- Coverage. Percentage of record categories with an approved period, trigger and owner.
- Automation. Percentage of categories where disposal runs automatically in at least the primary system.
- Overdue volume. Number of records past their retention period and not yet disposed of. This is the number that matters most.
- Hold hygiene. Number of active legal holds, and number older than 24 months without review.
- Vendor coverage. Percentage of vendors holding company data with a contractual deletion clause.
Overdue volume is uncomfortable the first time it is measured. It is also the metric that turns retention from a policy into a budget line, because it shows the size of the exposure in a unit the board understands.
The 30, 60, 90 day roadmap
Days 1 to 30: decide and inventory
- Appoint the owner and a working group with legal, IT, finance, HR and one commercial function.
- List record categories in business terms. Stop at sixty.
- Map each category to its primary system and known copies.
- Collect legal minimums and maximums for each jurisdiction, with sources.
- Identify the three categories with the largest volume and the weakest reason to keep.
Days 31 to 60: write and approve
- Draft the policy: principles, roles, exceptions, legal hold procedure.
- Draft the schedule with period, trigger, source, disposal method and systems.
- Check that every trigger exists as a date in the system. List the missing fields.
- Align the privacy notice with the schedule.
- Get executive approval. Publish.
Days 61 to 90: delete something
- Turn on native retention settings in email and collaboration tools.
- Run the first deletion on one high volume category, on a copy first, then in production.
- Start the disposal log.
- Test the legal hold procedure on a simulated dispute.
- Add retention and exit clauses to the vendor contract template.
- Report the five metrics for the first time.
The title of the third phase is deliberate. A policy that has not deleted anything in ninety days will not delete anything in a year. One real deletion, logged and reported, changes how the organization sees the whole program.
How retention connects to the rest of governance
Retention is one of several controls that share the same inventory. A data governance framework defines who owns the data. The retention schedule defines how long it lives. A business continuity plan defines how it is restored after an incident. When these three are written by different teams with different lists of systems, they contradict each other, usually on backups.
The practical rule is one inventory of systems and record categories, referenced by all three documents. It is less elegant than a single platform and much cheaper, and it can be built in a spreadsheet before any tool is bought.
In a sports distribution company I advised, sales grew 30 percent after we rebuilt marketing on customer data. The work started with a decision about which contacts were still customers and which were history. Deleting the second group made every campaign cheaper and every report more honest. Less data, better chosen, is a commercial advantage before it is a compliance one.
If you want an outside view on where your policy will break before a regulator or an attacker finds out, a consultation request with your current schedule attached is the fastest way to get one.
FAQ
What is a data retention policy?
A data retention policy is an internal document that states how long a company keeps each category of record, why, who owns the decision and how the record is disposed of at the end of the period. It includes a retention schedule, which is the table of categories, periods and triggers, and a procedure for exceptions such as legal holds. It covers personal data and business records, in applications, email, shared drives and backups.
How do you write a data retention policy for a business?
Follow eight steps in order: appoint an owner, inventory records by category, state the purpose of each, collect legal minimums and maximums, set a period and a trigger, define the disposal method, write the exceptions, then approve and schedule the review. Map every category to the systems where it lives. Without that map the policy cannot be applied. Plan ninety days, and run at least one real deletion before the end.
How long should a company keep data?
There is no single answer. Minimum periods come from tax, employment and sector rules. In the United States the IRS baseline for tax records is three years, employment tax records must be kept at least four years, payroll records three years and personnel records one year. Privacy laws set the other limit: personal data should not be kept longer than necessary for its purpose. Each record category needs its own period and source.
Is a data retention policy required by law?
In many cases, yes, directly or in practice. European and UK data protection law requires companies to justify how long they keep personal data, and the UK regulator states that a policy with standard retention periods is needed wherever possible. California requires businesses to disclose retention periods to consumers. Sector rules in finance and healthcare require specific records to be kept for fixed periods. A written policy is the simplest way to show compliance with all of them.
What is the difference between a data retention policy and a backup policy?
A backup policy exists to restore systems and data after a failure, and it defines how often copies are made and how long they rotate. A data retention policy decides how long information should exist at all. The two must be consistent: if backups are kept for years, deleted data survives in them. The usual solution is a short fixed backup cycle, and a rule that restored data falls under the schedule again.
What happens to the retention schedule during a lawsuit?
When litigation is pending or reasonably anticipated, the company must preserve relevant records, even if the schedule says they should be deleted. This is a legal hold. The policy should state who can order a hold, which records and systems it covers, how automated deletion is suspended, and who releases the hold when the matter ends. Holds should be reviewed periodically, so that they do not become permanent by neglect.
How does artificial intelligence change data retention?
It adds new record types and raises the stakes on old ones. Prompts, outputs, indexes and training data all need a retention period and a vendor commitment. Internal assistants can surface old files that nobody had opened in years, which turns forgotten archives into active risk. A company should clean and classify its data before connecting it to an assistant, and should check how quickly a deleted document disappears from the tool's index.