ARTICLE
4 October 2026

Robotic Training And Personal Data: Key Considerations Under The DPDP Act And EU Framework

AP
Argus Partners

Contributor

Argus Partners is a leading Indian law firm with offices in Mumbai, Delhi, Bengaluru and Kolkata. Innovative thought leadership and ability to build lasting relationships with all stakeholders are the key drivers of the Firm. The Firm has advised on some of the largest transactions in India across various industry sectors. The Firm also, regularly advises the boards of some of the biggest Indian corporations on governance matters. The lawyers of the Firm have been consistently regarded as the trusted advisors to its clients with a deep understanding of the relevant business domain, their business needs and regulatory nuances which enables them to clearly identify the risks involved and advise mitigation measures to protect their interests.
Argus Partners maintains offices in three major Indian cities - Mumbai, New Delhi, and Bengaluru - providing legal services across the country. The firm's strategic presence in these key metropolitan areas enables comprehensive coverage of India's primary business and commercial centers. Contact information and physical addresses are provided for each location to facilitate client communication and engagement.
European Union Corporate/Commercial Law

Robotic systems are increasingly trained on human-generated data. Depending on the training method, this may include video and audio recordings, movement trajectories, tele-operation inputs, force measurements, gait, location and other sensor data. This raises an important data-protection question: when does training a robot amount to processing personal data, and how does the answer change depending on the way in which that training is undertaken?

The position is particularly interesting when India is compared with the European Union. In India, the Digital Personal Data Protection Act, 2023 (“DPDP Act”) (here) principally regulates the processing of personal data used in the training exercise. In the EU, the General Data Protection Regulation (“GDPR”) (here) applies to the personal-data layer, while the EU Artificial Intelligence Act (“AI Act”) (here) and, in appropriate cases, the EU Machinery Regulation may separately regulate the resulting robotic system.

How robots are trained

Neither the DPDP Act nor the GDPR separately defines “robotic training”. Industry standards similarly tend to classify robots by function rather than by training methodology. ISO 8373:2021, for instance, contains the general vocabulary for robotics and distinguishes between different categories of robots (here).

For data-protection purposes, however, the more useful distinction is between the manner in which the system learns. Common techniques include human demonstration or imitation learning, tele-operation, kinesthetic teaching, training on video or multimodal datasets, simulation and synthetic data, and continuous learning after deployment.

This distinction matters because the privacy implications differ considerably. A robot trained entirely in simulation may process no personal data at all. A robot trained by observing workers may process their image, voice and movements. A tele-operated robot may capture both its operator and persons in the surrounding environment. A system that continues learning after deployment may create an ongoing stream of personal data long after the initial development phase has ended.

Position under the DPDP Act

When does training data become personal data?

Section 2(t) of the DPDP Act defines “personal data” as data about an individual who is identifiable by or in relation to such data. “Processing” is defined broadly under Section 2(x) and includes collection, recording, storage, use, combination and sharing. Accordingly, the fact that a developer is interested in teaching a robot how a person moves, rather than in identifying that person, does not itself take the data outside the Act.

Video and audio recordings will ordinarily constitute personal data where an individual can be identified. Less obvious datasets require a more contextual analysis. A pose trajectory, grip-force measurement or inertial sensor reading may not identify anyone by itself, but may become personal data where it is linked with a session ID, employee identifier, device record or shift roster. 

This is particularly important for multimodal robotic datasets. Facial information may be removed from video, while the remaining movement, device and session data continue to permit the same individual to be identified. Facial blurring or removal of names should therefore not automatically be treated as anonymisation. The relevant question is whether the person remains identifiable from the dataset as it is actually held and used, including by combining it with other information reasonably available to the organisation.

No separate sensitive-personal-data category 

Unlike the GDPR, the DPDP Act does not create a separate category of sensitive or special-category personal data. Facial images, gait information or behavioural data therefore do not automatically attract a separate statutory processing threshold merely because of their sensitivity. Their sensitivity can nevertheless remain relevant. Under Section 10, the Central Government may designate a Data Fiduciary as a Significant Data Fiduciary having regard to matters including the volume and sensitivity of personal data processed and the risks to Data Principals. Significant Data Fiduciaries are subject to additional requirements, including periodic data protection impact assessments and audits.

What is the lawful basis for training on employee data? 

Section 4 permits processing for a lawful purpose on the basis of consent or one of the “certain legitimate uses” specified in Section 7. The DPDP Act does not contain a general “legitimate interests” ground comparable to Article 6(1)(f) GDPR, nor does it recognise contractual necessity as a general standalone basis.

Section 7(i) permits certain processing for purposes of employment and related purposes. However, using employees to generate data for development of a general robotics model is not necessarily the same as processing data for administering the employment relationship. Attendance, workplace access and occupational safety data can readily be connected to employment. Continuous recording of a worker's movement, field of view and manipulation behaviour to improve a commercial robotics model is more difficult to characterise in the same way.

Where the model-training purpose is distinct from the employment relationship, consent under Section 6 may therefore provide the more appropriate basis. That consent must be free, specific, informed, unconditional and unambiguous. In practice, particularly in an employment setting, this means that refusal or withdrawal should not lead to adverse employment consequences.

The final Digital Personal Data Protection Rules, 2025 (“DPDP Rules”) (here) also require the notice accompanying a consent request to describe the personal data and the specified purposes for which it is processed.

Training data and employee-monitoring data may be the same data 

A peculiar feature of robotic training is that the same information may support very different purposes. Movement speed, force measurements, task sequences or accuracy data can teach a robot how to perform a task. The same data can also show how quickly an employee works, the frequency of errors or differences in physical performance. 

The technical dataset may therefore remain unchanged while its legal purpose changes. If data is collected for robotics training, it should not automatically be repurposed for productivity monitoring, performance assessment or disciplinary action. If such additional uses are intended, they should be separately identified and assessed. 

This also has a technical implication. Where practicable, robotics-training data should be segregated from HR and workforce-management systems, with access limited to persons involved in the development process. 

Third-party facilities and outsourced data collection

A company may outsource the collection of robotic-training data without outsourcing responsibility for the processing. Under Section 2(i), the Data Fiduciary is the person who determines the purpose and means of processing. Accordingly, if the robotics developer determines the training objective, the sensors to be used, the information to be recorded and the way in which it will be processed, it may remain the Data Fiduciary even where the physical data collection is performed by another entity.

The third-party employer may separately act as a Data Fiduciary for its own payroll, attendance and employment-related processing. Accordingly, the allocation should follow the actual decision-making arrangement rather than the labels used in the agreement. Where the vendor processes training data on behalf of the developer, the contract should address processing instructions, security, sub-processing, assistance with Data Principal requests, retention and deletion.

Retention, erasure and the trained model 

Robotic training also raises a harder question: how long can the training data be retained, and what happens once it has been used to train a model?

Section 8(7) of the DPDP Act requires a Data Fiduciary to erase personal data upon withdrawal of consent or when it is reasonable to assume that the specified purpose is no longer being served, unless retention is required by law. Section 12 separately gives Data Principals a right to seek erasure, subject to continued retention being necessary for the specified purpose or compliance with law. This makes indefinite retention of identifiable raw training data difficult to justify merely on the basis that it may prove useful for future model development.

A more defensible architecture would distinguish between:

  1. raw identifiable recordings;
  2. pseudonymised or de-linked training datasets;
  3. genuinely anonymised derived data; and
  4. the trained model itself. 

Raw footage may need to be retained for a shorter period than derived movement or action data. Where identifying information is no longer required, it can be removed earlier in the pipeline. An additional issue arises where a Data Principal requests erasure after their data has already contributed to model training. The DPDP Act does not presently prescribe a general obligation to retrain a model or perform “machine unlearning” following every erasure request.

However, it should also not be assumed that model weights are necessarily outside data-protection law in every case. The more relevant question is whether information relating to individuals can still be recovered or inferred from the model.

The EDPB has taken a similar approach under GDPR. In Opinion 28/2024 (here), it rejected the assumption that AI models trained on personal data are automatically anonymous. Anonymity must be assessed case by case, including whether individuals can be identified from the model or whether personal data can be extracted through queries. This is likely to become particularly important for robotics foundation models trained on large amounts of human movement or audiovisual data. 

Publicly available training datasets

Section 3(c)(ii) of the DPDP Act excludes certain personal data from its application where that data has been made publicly available by the Data Principal herself or by another person who is legally required to make it publicly available. This could be relevant where a robotics developer trains systems on publicly posted videos of human activity.

The exclusion should nevertheless be applied carefully. Information being technically accessible online is not necessarily the same as the Data Principal having made it publicly available. Leaked material, unauthorised reposts or data obtained from restricted sources may fall outside the exclusion. This is also a significant point of divergence from Europe: personal data does not generally fall outside GDPR merely because it is publicly accessible. 

Cross-border processing 

Robotics development frequently involves centralised global training infrastructure, meaning data collected in one jurisdiction may be transferred to engineering or research teams elsewhere. Section 16 of the DPDP Act permits the Central Government to restrict transfers to specified countries or territories. The Indian framework therefore does not adopt the GDPR model of adequacy decisions, Standard Contractual Clauses or other Chapter V transfer mechanisms as the basic statutory structure. 

From an operational perspective, however, contracts with overseas recipients should still address purpose restrictions, security, further disclosure, deletion and assistance with Data Principal rights, particularly because the Data Fiduciary remains responsible for processing undertaken on its behalf. 

Position under the GDPR 

The GDPR begins from a similar concept of personal data but differs materially in its lawful-basis and sensitive-data architecture. Articles 4 and 5 apply to identifiable video, movement, behavioural and sensor data even where the objective is to train a machine rather than identify an individual.

Article 6, however, provides a wider range of lawful bases. In particular, legitimate interests may potentially support AI development where the controller establishes the interest pursued, demonstrates necessity and shows that it is not overridden by the rights and interests of affected individuals. 

The EDPB's Opinion 28/2024 (here). applies this familiar three-stage analysis specifically to AI models and also considers the individual's reasonable expectations, including the source and context in which the underlying information was collected. This makes the EU position more flexible in one respect than the DPDP Act, although that flexibility is subject to a proportionality and balancing analysis. 

Employee consent is also treated particularly cautiously under GDPR because employees may not be in a position to refuse freely. Depending on the Member State, Article 88 and domestic employment privacy rules may impose additional requirements.

Biometric data under GDPR 

Article 9 creates a further distinction absent from the DPDP Act. Biometric data processed for the purpose of uniquely identifying an individual constitutes special-category personal data and ordinarily requires both an Article 6 lawful basis and satisfaction of an Article 9 condition. The purpose qualification is important. Ordinary footage does not become Article 9 biometric data merely because a person's face appears in it. Similarly, movement or gait information does not automatically become special-category data. 

The position changes where technical processing is undertaken to create facial templates, gait signatures, voiceprints or similar identifiers for the purpose of uniquely recognising individuals. 

Bystanders and incidental collection 

Human participants are not necessarily the only Data Principals affected by robotic training. A tele-operated robot, body-mounted camera or continuously learning system may also capture people in the surrounding environment. Those individuals may have no contractual or employment relationship with the developer and may not even know that their data is entering a training dataset. 

This makes certain training environments materially more difficult to justify than controlled data collection. Under GDPR, large-scale systematic monitoring of publicly accessible areas is one of the express circumstances identified in Article 35 as potentially requiring a Data Protection Impact Assessment (“DPIA”). The DPDP Act takes a different approach: the express DPIA obligation arises for Significant Data Fiduciaries under Section 10 rather than automatically because a particular processing operation crosses a high-risk threshold. 

The additional EU AI Act layer 

The GDPR regulates the data used for robotic training. The AI Act may separately regulate the system that results from that training. Not every AI-enabled robot is a high-risk AI system. Under Article 6, a system may be classified as high-risk either because it forms part of the regulated product-safety framework or because its intended purpose falls within one of the use cases in Annex III. For robotics, this means that classification depends on what the AI actually does. 

An AI system forming a safety component of regulated machinery may enter the product-related high-risk regime. A robotic system used to make certain decisions about workers may instead fall within the employment-related categories in Annex III.

For relevant high-risk systems using data-based training techniques, Article 10 imposes requirements around training, validation and testing datasets. These include the origin and collection of data, preparation and labelling, representativeness, possible bias, gaps and shortcomings and, notably, the original purpose for which personal data was collected. That point is important for robotics developers considering reuse of workplace, CCTV or operational datasets. Under the European framework, the provenance of the dataset may therefore matter under both GDPR and the AI Act.

The AI Act also goes further than general data protection law by prohibiting certain practices. These include, subject to its detailed conditions, untargeted scraping of facial images from the internet or CCTV to create or expand facial-recognition databases and certain uses of emotion-recognition systems in workplaces. Accordingly, having a lawful basis under GDPR does not necessarily mean that the intended AI use is lawful under the AI Act.

Continuous learning and the Machinery Regulation 

There is also an important difference between a robot that is trained before deployment and one that continues to learn afterwards. Where the model is trained offline and effectively fixed before deployment, the data-protection issue is principally concentrated in the development phase. A continually learning robot creates an ongoing processing activity. New interactions may continuously enter the model-training pipeline, potentially involving employees, customers or bystanders. 

In Europe, this may also engage product-safety concerns. The Machinery Regulation expressly contemplates machinery with fully or partially self-evolving behaviour or logic and requires lifecycle risks arising from such behaviour to be considered. Continuous learning therefore changes more than the privacy analysis: the resulting change in physical behaviour may itself become a safety-regulatory issue. 

How the legal position changes with the training method 

Training method

Principal issue

Simulation / synthetic training

Ordinarily outside DPDP/GDPR where no identifiable human data is used, although the resulting system may still be regulated under EU AI/product law

Human demonstration / imitation learning

Video, movement and sensor data may be personal data; employee lawful basis and purpose limitation become important

Tele-operation

Both the operator and persons appearing in the robot's environment may become Data Principals

Kinesthetic teaching

Force and movement data may be personal where retained with demonstrator identifiers; risk reduces where data is genuinely de-linked

Body-mounted / egocentric recording

Particularly likely to capture third parties and surrounding contextual information

Internet or third-party datasets

Requires assessment of provenance; India's public-data exclusion differs materially from GDPR

Biometric identification training

GDPR Article 9 may apply; DPDP has no equivalent special-category regime

Continuous learning after deployment

Creates continuing privacy obligations and, in the EU, may additionally engage AI and machinery regulation

Genuinely anonymised derived data

Outside DPDP/GDPR where individuals can no longer reasonably be identified

Where India and the EU ultimately differ

The main difference is structural rather than simply one of regulatory strictness. Under the DPDP Act, the analysis remains centred on the processing of digital personal data: what data is being processed, whether the individual is identifiable, the specified purpose, the applicable basis for processing, and when the data must be erased. 

The European position may require several additional layers of analysis. GDPR governs the personal data and provides a wider set of lawful bases. Article 9 creates a separate regime for special-category data. The AI Act may regulate training-data governance and the intended use of the resulting system. The Machinery Regulation may become relevant where machine learning affects the behaviour and safety of the physical product.

At the same time, India contains a broader exclusion for certain publicly available personal data, while the EU does not generally remove such data from GDPR merely because it can be accessed online.

Conclusion

The data-protection implications of robotic training depend principally on how the system is trained and whether identifiable human data enters or remains within the training pipeline. Simulation and genuinely anonymised datasets may fall outside the DPDP Act and GDPR, while human demonstration, tele-operation and continuous learning can engage more substantial obligations relating to lawful basis, purpose limitation, retention and data-subject rights. Accordingly, robotics developers should assess personal-data exposure across the training lifecycle and, where possible, minimise or remove identifying information at an early stage. In the EU, this analysis may additionally need to account for the AI Act and Machinery Regulation, depending on the intended use and operation of the robotic system.

This update has been contributed by Jitendra Soni (Partner), Sadia Akhter (Senior Associate), and Aditya Kashyap (Associate).

The content of this article is intended to provide a general guide to the subject matter. Specialist advice should be sought about your specific circumstances.

[View Source]

Mondaq uses cookies on this website. By using our website you agree to our use of cookies as set out in our Privacy Policy.

Learn More