BlogGEO

The EDPB’s Draft AI Web-Scraping Guidance: Why Form Teams Should Not Treat Submissions as AI Training Data by Default

The EDPB’s draft guidance on web scraping for generative AI is aimed at AI development, but it gives form teams a timely reason to separate operational use of submissions from model training and fine-tuning.

Four lanes for AI form data: build the form, process submissions, enrich records, and train models.
Thimo Waanders
Thimo Waanders
Founder & Lead Funnel StrategistUpdated

On 7 July 2026, the European Data Protection Board adopted draft Guidelines 03/2026 on web scraping in the context of generative AI. The EDPB announced the draft and opened its public consultation on 8 July, with comments due by 30 October 2026. The draft does not ban AI or web scraping. It does make casual repurposing of personal data for generative-AI development harder to justify.

The guidance directly concerns personal data scraped for generative-AI development or fine-tuning. It addresses lawful basis, transparency, data minimisation, special-category data, data-subject rights, and safeguards across the AI lifecycle. Read the EDPB announcement and the draft Guidelines 03/2026.

For marketing, sales, research, and product teams, the operational implication reaches beyond web crawling. Someone who submits a demo request, application, survey response, or intake form may expect follow-up, routing, qualification, and perhaps limited verification. That does not mean their response should automatically become a reusable corpus for training or fine-tuning a model.

This article offers an operating framework, not legal advice. The right assessment depends on the data, jurisdiction, notices, lawful basis, contracts, vendors, and safeguards in a specific workflow.

What the EDPB actually said about AI web scraping

Web scraping is the automated collection of information from websites or other online sources. The EDPB’s draft focuses on personal data scraped for developing or fine-tuning generative-AI models. Its scope includes organisations that scrape data themselves, use a third party to scrape, or obtain already-scraped datasets for those purposes. The draft guidelines describe those scenarios.

The important correction is straightforward: public availability is not unrestricted reuse. GDPR continues to apply when scraped information contains personal data. Making data public online does not itself provide consent for a different purpose such as AI training. Nor is the absence of a robots.txt file consent to scrape personal data for that purpose. Technical opposition signals and access restrictions can be relevant factors, but they are not a complete compliance answer.

The draft also stresses minimisation before collection. It points to precise criteria, data mapping, filters, and consideration of synthetic or pseudonymised alternatives. It identifies greater caution for sources aimed mainly at minors or likely to contain sensitive information. Where special-category personal data is involved, both an Article 6 lawful basis and an Article 9 condition or derogation are required. The EDPB’s release summarises these themes.

These are draft guidelines, not new law, a final decision, or an enforcement action. They are, however, the EDPB’s current interpretation of GDPR obligations in this context.

Why this matters if your team does not crawl the web

Your team may never operate a crawler. You may still use a vendor that performs public-profile research, relies on already-scraped data, retains prompts or uploaded files, provides AI-based enrichment, or offers model-improvement settings. The draft is relevant to this wider chain because it covers third-party scraping and already-scraped datasets used for generative-AI development or fine-tuning.

That does not make enrichment the same as model training. They are different activities. The operational task is to identify each purpose and recipient instead of calling every AI-enabled action “automation.” Reed Smith’s analysis of the draft also highlights the importance of source controls, reasonable expectations, rights, and lifecycle safeguards.

Start with five questions:

  1. What did the visitor provide directly?
  2. What do we add through enrichment or verification?
  3. Which systems receive data after submission?
  4. What does each system do with it: route, classify, verify, analyse, retain, or train?
  5. Could we explain that path in language that matches the visitor’s experience?

If the answer to the last question is no, the problem is usually not a missing checkbox. It is an unclear purpose or an uncontrolled destination.

A form submission is not blanket permission for every AI use

Consider a B2B demo request. A visitor shares their name, work email, company, team size, and a short description of their problem. The apparent operational purpose may be to assess fit, route the request, contact the visitor, and prepare a relevant conversation.

Several distinct activities may follow:

  • Collection: capture the answers needed to handle the request.
  • Operational processing: classify the request, assign an owner, create a record, or send a reply.
  • Enrichment or verification: use an email address, domain, or company website to fill missing business context or check an address.
  • Model training or fine-tuning: add raw requests, attached files, applications, or sales context to a dataset used to improve a model.

The first three may fit within a defined lead-handling workflow, subject to appropriate review. The fourth changes the purpose and should have its own privacy, legal, security, and vendor assessment. It should not be treated as an incidental benefit of accumulating useful responses.

Set that boundary before data reaches a training environment. Removing personal data after it has been incorporated into a trained model can be difficult, as noted in independent legal analysis of the draft.

The four-lane AI data audit for forms and funnels

This is Stepform editorial guidance informed by the EDPB’s themes. It is not EDPB terminology. Its purpose is to stop teams from collapsing different processing activities into one vague “AI use” label.

AI data lanePractical exampleOperational treatmentReview trigger
1. AI builds the formGenerate a first draft, rewrite a question, or draft conditional logic.Keep a human owner responsible for wording, field mapping, routing, and publishing.Review before publishing, especially when prompts include real customer or prospect data.
2. AI processes the submissionClassify inquiry type, summarise free text, or route a request.Document the purpose, inputs, recipients, access, and fallback for incorrect output.Escalate when decisions materially affect a person or unstructured inputs may contain sensitive context.
3. AI enriches or verifies the recordVerify an email address or fill missing company details from a business domain.Limit inputs to what is needed. Identify the provider, source, and added data.Review sources, retention, rights support, inferred fields, and whether enrichment is necessary.
4. AI trains or fine-tunes a modelUse applications, uploads, survey responses, or sales notes to improve a model.Treat this as a separate project with a defined dataset, purpose, owner, safeguards, and deletion approach.Require separate privacy, legal, security, and vendor review before inclusion.

The practical principle is purpose limitation in operational terms: do not move data from one lane to another because it is technically convenient. The EDPB draft discusses pre-collection filters, sensitive-data handling, data-subject rights, and lifecycle safeguards. Read the primary guidance for its full context.

Build an AI-aware lead-capture flow

A useful form gives the visitor a clear exchange and gives the team only the context it needs. Here is a practical B2B demo-request flow.

Page 1: State the exchange

Ask for a work email and company website. Explain the next step in specific language, such as: “We’ll use your details to review your request and follow up about a relevant next step.” This is illustrative product copy, not legal boilerplate.

Page 2: Qualify only when the answer changes the next step

Ask role, team size, primary use case, and timeline. Use a branch when an answer changes routing or the conversation. For example, route a support request to support instead of asking that person for sales budget. Do not collect budget, internal tools, or open-ended background because it might be useful someday.

Page 3: Explain purposeful additions

If you verify a work email or enrich a company record for qualification, explain that near collection in plain language. Ensure the live workflow matches the explanation. A company-domain lookup should not silently become unrelated public-web research or model training.

After submission: map, route, and preserve boundaries

Map answers into structured Person, Company, and custom fields. Preserve UTM and referral context separately from visitor answers. Route qualified requests to an owner, send a confirmation, and assign a pipeline stage. Keep raw free text, files, audio, video, and notes out of training datasets unless a separate reviewed process authorises that use.

This is where Stepform becomes relevant. It supports AI-assisted drafting with human review, visual conditional logic, hidden fields and UTM capture, structured Person and Company records, enrichment or email verification when needed, pipeline management, and automations. The value is a visible managed-form workflow, not a claim that the product makes a workflow legally compliant.

Questions to ask every AI or enrichment vendor

The EDPB’s draft makes vendor and dataset provenance more important where an organisation obtains already-scraped data or uses third parties. These questions help operators collect facts for privacy, legal, security, and procurement review.

Ask the vendorWhy it mattersWhat a useful answer clarifies
Do you use our inputs, outputs, files, or metadata to train or fine-tune models?Operational processing and model improvement are distinct purposes.Whether training occurs, which settings control it, and whether the choice applies across product tiers or API routes.
How long do you retain submitted data, prompts, outputs, and logs?Retention should align with defined purpose and minimisation decisions.Retention periods, deletion processes, backup treatment, and access controls.
What data do you receive, and which subprocessors receive it?A form owner needs a destination map, not a generic security statement.Input fields, derived fields, regions, subprocessors, and onward transfers.
What are the sources for enrichment or verification data?The draft covers direct scraping, third-party scraping, and already-scraped datasets for AI development.Source categories, collection practices, available documentation, and public-web inputs.
How do you support access, deletion, correction, and objection requests?The EDPB treats rights mechanisms as relevant safeguards.Request channels, timelines, scope, and limitations across systems.
Do you create sensitive or inferred attributes?Sensitive information and inferences can create higher-governance questions.Generated fields, confidence indicators, controls to disable them, and error-correction options.

A vendor’s claim that it is “GDPR-ready” is not a complete answer. Your team still needs to understand the actual data path, purposes, settings, and contractual roles. In large-scale scraping scenarios where an individual-notice exception may apply, the EDPB says a public notice should still explain the processing, data categories, purposes, legal basis, and source information as far as feasible. See the guidelines for the relevant conditions.

Common mistakes in AI form-data governance

  • Using a vague AI notice. “We may use AI to improve our services” does not explain the operational path. Describe the purpose, recipients, and consequences in reviewable language.
  • Assuming public profiles are reusable. Public availability does not remove GDPR obligations where personal data is involved. Publishing data online is not consent for a new AI-training purpose.
  • Collecting for hypothetical future use. A field without a routing, qualification, or service purpose is a minimisation problem before it becomes an AI problem.
  • Mixing routing and training in one setting. A vendor can process a response for your workflow while a separate setting uses it for model improvement. Audit both.
  • Ignoring unstructured responses. Free text, uploads, audio, video, and applications can contain more context than fixed-choice fields. Give them tighter access, destination, retention, and training boundaries.
  • Treating consent as the only governance tool. The EDPB indicates consent will usually not be workable for large-scale external scraping because the scraper often lacks a direct relationship with affected people. That does not resolve every first-party workflow, where lawful basis and transparency remain context-specific. The draft explains the legal-basis analysis.

A 30-minute audit for an existing form workflow

  1. List the forms. Include lead forms, applications, quote requests, surveys, waitlists, support intake, and calendar pre-qualification flows.
  2. Write the promise. Record what each visitor is told will happen after submission.
  3. Map every destination. Include CRM records, spreadsheets, inboxes, Slack, verification services, enrichment providers, AI tools, analytics, webhooks, and archives.
  4. Label each destination by lane. Is it form building, operational processing, enrichment or verification, or model training and fine-tuning?
  5. Remove unclear fields and destinations. If a field or recipient lacks a current purpose, remove it or stop sending data there.
  6. Test the experience. Submit a test response and compare its actual path with the explanation on the form. Include partial-submission handling in the test.

FAQ

Did the EDPB ban AI web scraping?

No. The EDPB adopted draft guidance, not a blanket ban or a final enforcement decision. It adopted the draft on 7 July 2026, announced it and opened consultation on 8 July, and set 30 October 2026 as the deadline for comments. The draft explains how GDPR applies when personal data is scraped for generative-AI development or fine-tuning. <a href="https://www.edpb.europa.eu/news/news/2026/edpb-sheds-light-anonymisation-and-web-scraping-generative-ai-and-adopts-final-version-guidelines-blockchain_en">Read the EDPB announcement</a>.

Does the guidance say we cannot use form submissions with AI?

No. The guidance directly concerns web scraping for generative-AI development or fine-tuning. For first-party forms, its practical relevance is to document and distinguish collection, follow-up, enrichment, automation, analytics, and any separate training use. <a href="https://www.edpb.europa.eu/our-work-tools/our-documents/guidelines/guidelines-032026-web-scraping-context-generative-ai_en">See the draft’s scope</a>.

Can we enrich a lead record with company information?

Enrichment is distinct from model training and may be useful for qualification or routing. It still needs a purpose-specific assessment. Identify the provider, inputs, data sources, retention, recipients, rights process, and whether enrichment is necessary for the workflow.

Does a consent checkbox make AI training compliant?

Not automatically. The EDPB says publishing personal data online does not itself give consent for a different purpose such as AI training. It also indicates consent is often not workable for large-scale external scraping. The appropriate approach depends on the full facts and should be reviewed with counsel. <a href="https://www.edpb.europa.eu/our-work-tools/our-documents/guidelines/guidelines-032026-web-scraping-context-generative-ai_en">Read the draft guidance</a>.

Are public LinkedIn profiles safe to use for AI training?

Do not assume so. Public availability does not remove GDPR obligations where scraped information contains personal data. The EDPB draft calls for purpose-specific analysis of lawful basis, transparency, minimisation, expectations, safeguards, and rights.

How should we handle uploads, video, voice, and free-text answers?

Treat them as higher-context inputs. Limit collection to what the workflow needs, control access and destinations, set retention deliberately, and do not add them to training or fine-tuning datasets by default. If sensitive information may be present, seek specialist privacy and legal review.

Sources

Build a managed form workflow, not an ungoverned data collection point

The EDPB’s draft is a timely reminder that data is not interchangeable because it is accessible or useful. A strong form workflow makes purpose visible: what is collected, why it is collected, where it goes, what is added, who acts on it, and what cannot happen without separate approval.

For form teams, that means minimised fields, purposeful branches, structured records, reviewable enrichment, clear routing, and a firm boundary between workflow assistance and model training. Stepform helps teams design and manage workflows across forms, submissions, pipelines, automations, and analytics. It does not replace a privacy assessment, vendor due diligence, contractual analysis, or legal review.

More articles

Bright orange and blue abstract shapes against sky
Lead generation

ChatGPT Ads Need a Lead-Quality Loop, Not Just a Conversion Pixel

CallTrackingMetrics’ new OpenAI Ads attribution integration puts qualified-conversion feedback in focus. Here is how to design a ChatGPT Ads measurement loop that separates clicks and form fills from sales-ready demand.

Thimo Waanders
Thimo Waanders
Founder & Lead Funnel Strategist
Client intake funnel for agencies with project fit, scope, budget, files, and follow-up
Funnels

How to Create a Client Intake Funnel for Agencies and Service Businesses

Learn how to create a client intake funnel that collects the right context, qualifies fit, handles files, routes requests, and keeps new inquiries organized.

Thimo Waanders
Thimo Waanders
Founder & Lead Funnel Strategist
B2B lead workflow from anonymous visit through a qualification form, structured records, routing, and a reviewed AI-generated account summary.
GEO

6sense MCP Server Open Beta: Why AI-Accessible Intent Makes Governed Lead-Form Data More Important

6sense’s MCP Server open beta puts selected GTM intelligence into compatible AI-agent environments. Learn why declared form answers, attribution, structured records, explicit routing, and human review still matter.

Thimo Waanders
Thimo Waanders
Founder & Lead Funnel Strategist