What Is a Data Processing Agreement for EdTech AI?
A technical breakdown of the clauses EdTech AI vendor contracts need beyond a standard DPA template, covering FERPA, COPPA, model training limits, retention, and subprocessor disclosure.
What an EdTech AI DPA Must Cover
Beyond baseline processor terms, institutions should negotiate explicit coverage in these areas:
Model Training Limits
Whether student data may be used to train or fine-tune AI models, and under what consent basis.
Retention and Deletion
Separate terms for input data, model-derived artifacts, and AI-generated outputs.
Subprocessor Disclosure
A maintained schedule identifying downstream vendors and chained foundation models.
Audit Rights
Contractual access to logs verifying runtime processing matches stated terms.
Why Generic Vendor DPAs Fall Short for AI-Driven EdTech
A standard data processing agreement typically defines the subject matter, duration, categories of data, and processor obligations for a vendor relationship, structured around static activities like storage, transmission, and access control. GDPR Article 28 reflects this baseline, requiring processors to be authorized before engaging subprocessors and to process data only on documented instructions. This structure assumes data processing is a fixed, observable transaction.
AI systems break that assumption. Training and inference are distinct processing activities, and permission for one does not imply permission for the other. AI systems may retain student input as part of model context windows, fine-tuning datasets, or logs in ways a static DPA never anticipates. Vendor stacks frequently chain multiple third-party foundation models or inference providers, creating subprocessor layers invisible in a traditional single-vendor data flow diagram. Outputs such as tutoring responses or assessment scores can themselves contain or reconstruct student information, requiring retention and deletion terms distinct from the original input data. A DPA drafted for legacy, single-region EdTech systems will not address any of this without AI-specific clauses added deliberately.
Key distinction: Training and inference are separate processing activities. Authorization for inference alone should not be read as permission to train or fine-tune on student data.
Regulatory Foundations: FERPA, COPPA, and State Student Privacy Laws
FERPA's school-official exception under 34 CFR Part 99 permits disclosure of education records to a vendor only if the vendor operates under the direct control of the school regarding use and maintenance of those records, and only for the authorized educational purpose. FERPA does not explicitly define AI model training as a distinct category of use, which leaves interpretation of "authorized purpose" to negotiation between institution and vendor.
COPPA, under 16 CFR Part 312, requires verifiable parental consent to collect personal information from children under 13, though schools may consent on behalf of parents solely for use within an authorized educational context. The FTC's May 2022 policy statement on education technology clarified that providers cannot use children's data collected for educational purposes for other commercial purposes, including product or model improvement, without separate parental consent. This directly constrains whether a DPA can permit an AI vendor to use student data for model training at all.
State laws add further restriction. California's SOPIPA (Cal. Bus. & Prof. Code § 22584) prohibits school service providers from using student data to build non-educational profiles, run targeted advertising, or sell student information. Because state statutes vary in scope and enforcement, no single DPA template can be assumed compliant across all jurisdictions. Department of Education PTAC guidance recommends contracts explicitly define permitted data uses, security safeguards, and destruction or return requirements upon termination, and NIST's AI RMF offers a reference structure for documenting data provenance and lifecycle management where an institution wants AI-specific documentation beyond what FERPA or COPPA require directly.
- Define permitted educational uses narrowly enough that model training is either explicitly allowed or clearly excluded.
- Align COPPA consent scope with any commercial or product-improvement use of student data.
- Account for state statutes (such as SOPIPA) that restrict profiling, advertising, and sale of student information.
- Document destruction or return requirements at contract end, including AI-derived artifacts where applicable.
Structuring Subprocessor and AI Model Chaining Disclosures
Traditional data flow diagrams assume a single vendor processes data end to end. AI vendor stacks rarely work this way. A tutoring or analytics platform may route student input through one or more third-party foundation models, separate inference providers, and logging or monitoring layers, each of which may retain or process data differently. A DPA that only names the primary vendor leaves these downstream relationships unexamined.
An EdTech AI DPA should require the vendor to maintain a current, accessible schedule of subprocessors that explicitly includes any third-party AI models used for inference or training support, not just conventional infrastructure providers. Following the structure in GDPR Article 28, the agreement should require prior authorization or advance notice before a new subprocessor or model provider is added, and it should require disclosure when the underlying model version or provider changes, since different foundation models can carry different default retention or training behaviors. Without this, an institution may approve a vendor relationship without ever knowing which AI systems actually touch student records.
Practical requirement: Treat foundation model providers and inference vendors as subprocessors. Require notice when the model version or provider changes, not only when a new legal entity is added.
Audit, Logging, and Enforcement Provisions
Contractual commitments describe intended behavior; they do not confirm actual behavior. A DPA that prohibits training on student data or sets a defined retention period is only as reliable as the institution's ability to verify compliance. This is an emerging area of practice, and no standardized audit methodology for AI vendor compliance currently exists, but the DPA should still include specific mechanisms rather than relying on periodic vendor attestations.
Practical provisions include the right to review processing and access logs, evidence of data deletion at the end of a retention period, and documentation of which AI models processed which categories of data. Runtime governance and audit logging capabilities, the kind covered under AI governance and runtime monitoring tooling, can support this verification by providing institutions with independent evidence of how student data was actually processed at runtime, which is useful both for ongoing vendor oversight and for demonstrating FERPA and COPPA due diligence if a regulator asks how compliance was confirmed rather than assumed.
Key AI-Specific Clauses to Negotiate
When reviewing or drafting an EdTech AI DPA, prioritize language that closes the gaps generic templates leave open:
- Explicit prohibition or tightly scoped permission for training and fine-tuning on student data.
- Separate retention, deletion, and return terms for inputs, logs, embeddings, and model-generated outputs.
- A living subprocessor schedule that includes chained AI models and inference providers.
- Advance notice or authorization rights before new model providers or versions are introduced.
- Audit and logging rights sufficient to verify runtime processing against contractual limits.
- Clear mapping of permitted uses to FERPA school-official purpose and COPPA educational context.
Verify AI Vendor Compliance at Runtime
A DPA sets the contractual commitment. Runtime governance and audit logging help confirm that an AI vendor's actual data handling matches it.
Explore Runtime Governance