
The fact that personal information is publicly accessible does not, by itself, resolve the legal consequences of collecting, combining and reusing it. This distinction is particularly important under the Digital Personal Data Protection Act, 2023 (“DPDP Act”), which expressly excludes certain publicly available personal data from its scope. At the same time, the rapid growth of automated scraping, data aggregation and artificial intelligence (“AI”) raises new questions about how that exclusion should apply to the large scale collection and repurposing of publicly available personal information.
The statutory exclusion
Section 3(c)(ii) of the DPDP Act excludes from its application personal data made or caused to be made publicly available by the Data Principal to whom the personal data relates, or by another person who is legally required to make such personal data publicly available.
The exclusion is therefore not framed simply by reference to whether information is accessible online. It turns on how and by whom the personal data was made or caused to be made publicly available. A professional may publish their name, designation and biography on a company website, or maintain a public social-media profile. Corporate or regulatory law may also require certain personal information to be publicly disclosed. These circumstances are distinct from information obtained through unauthorised access or from a restricted database.
The more difficult issue is what happens after information has entered the public domain. The statutory exclusion creates a boundary between public accessibility and subsequent exploitation. Whether the exclusion extends to every subsequent use of such data, including aggregation, enrichment and profiling, is likely to become an important question of statutory interpretation and will become increasingly significant as organisations prepare for the substantive operation of the DPDP framework.
Scraping changes the scale
Web scraping sits at this intersection. It automates the extraction of information from websites and other publicly accessible sources. Scraping supports legitimate activities including search indexing, price comparison, market research and academic research. It is also used to create lead-generation databases, data-broker products, training datasets and AI systems.
The legal issue is not whether scraping is inherently permissible, but what follows from it.
An organisation compiling publicly available professional information for an internal directory presents a different problem from one that scrapes millions of profiles, combines them with information from other sources, enriches the resulting dataset and sells commercial or behavioural insights.
This does not make every instance of scraping public personal data unlawful. The point is narrower: the scale, purpose, context and consequences of subsequent processing may materially alter the legal character of the activity.
The distinction is particularly important where information that was individually innocuous becomes significantly more revealing once aggregated. Each item may have been publicly observable at its source, while the resulting dataset may reveal considerably more about an individual than any individual source did.
Aggregation and inference
Aggregation sharpens the difficulty because privacy risks often arise not from individual facts, but from the ability to connect them.
A person’s employer, city, professional history, and social-media activity may independently be public. Combined, however, these facts can facilitate inferences concerning income, lifestyle, relationships, interests or purchasing behaviour that the individual never disclosed.
This also distinguishes observation from inference. A public webpage may disclose that an individual works for a particular organisation, a database may combine that information with other material to infer seniority, purchasing power or likely interests. Those inferences may never themselves have been made public by the individual.
The statutory exclusion therefore raises a difficult issue of statutory interpretation: whether publicly available information that is systematically aggregated and transformed into a substantially more detailed profile remains within the scope of the exclusion.
The DPDP Act does not expressly establish a doctrine of contextual integrity, nor does it provide a detailed framework for determining when aggregation changes the legal status of public information. That uncertainty is likely to become more significant as data-enrichment businesses and people-search products expand.
AI and the scale problem
AI sharpens the issue further because developers can collect vast quantities of publicly accessible material for training, evaluation, retrieval and model development.
The scale is fundamentally different from conventional human access. The relevant question is not merely whether an AI developer can access a webpage, but whether millions or billions of public data points can be systematically harvested, retained, combined and used to develop a commercial product, particularly where the underlying material contains personal data.
A public social-media post may have been deliberately placed in the public domain. Its incorporation into a large training corpus, however, places it in a different informational environment. It may subsequently contribute to model training, generate inferences or potentially be reproduced through an AI system.
This does not, by itself, establish that such processing is unlawful. Rather, it illustrates why the meaning and limits of the statutory exclusion will become increasingly important in an AI-driven information economy.
Context still matters
Social media provides a clear illustration. A public LinkedIn profile may be intended to make an individual professionally discoverable. A public Instagram account may be maintained for social visibility. That does not necessarily resolve whether systematic harvesting, aggregation with external datasets and commercial exploitation fall within the statutory exclusion.
The same distinction arises with public records. A court judgment may contain personal information because principles of open justice require its publication. Corporate filings may contain personal information because company law mandates disclosure. The purpose of the original disclosure is transparency or legal compliance, and subsequent commercial aggregation may serve an entirely different purpose.
This does not mean that every secondary use of publicly available information is prohibited. It does mean that public disclosure and unrestricted reuse are conceptually different propositions.
A question of purpose, not merely visibility
The future interpretation of Section 3(c)(ii) will require Indian law to distinguish accessibility from subsequent exploitation.
An interpretation under which every piece of publicly available personal data remains subject to the full DPDP framework would risk undermining the express statutory exclusion. Conversely, treating public availability as an unrestricted licence for indefinite collection, aggregation, profiling and commercial exploitation could facilitate the reconstruction of detailed personal profiles from information that individuals disclosed only in fragments.
The more appropriate inquiry is likely to be contextual: what information was made public, by whom and in what circumstances; what is the purpose of the subsequent processing; has the information been aggregated or enriched; does the resulting activity create materially greater risks for individuals; and how closely is the subsequent use connected to the character of the original disclosure?
These questions will become increasingly important as scraping, data brokers and AI systems transform the economics of publicly available information.
The DPDP Act recognises a meaningful distinction between information kept private and personal data deliberately placed in the public domain. It has not, however, resolved every consequence of that distinction.
The central challenge for Indian data-protection law will be determining whether, and to what extent, the statutory exclusion extends to subsequent collection, aggregation, inference and commercial exploitation.
In the age of automated scraping and AI, privacy may be affected not through a single disclosure, but through the accumulation and recombination of individually public facts. The law will increasingly have to confront that difference.













