#FactCheck- AI-generated videos and fake claims targeting Delhi CJP protests circulate on social media.
Research Wing
Innovation and Research
PUBLISHED ON
Aug 1, 2026
10
Executive Summary
Even after the conclusion of the Cockroach Janata Party (CJP) protest at Delhi’s Jantar Mantar, several misleading claims related to the demonstration continue to circulate on social media. During the protests, multiple AI-generated videos, deepfakes and fabricated narratives were circulated with the intention of misleading users. CyberPeace Research Wing’s Fact Check Team investigated several such viral claims and found that many of the videos were digitally manipulated or created using artificial intelligence (AI). The research revealed that fake audio, misleading captions and altered visuals were used to falsely attribute statements to public figures, including police officials, military officers and Union ministers. The following are four major false claims debunked by the CyberPeace Research Wing:
Claim 1: Delhi Police DCP Sumit Jha resigned from IPS in support of student protests
A video of Delhi Police DCP Sumit Jha was widely shared on social media with the claim that he resigned from the Indian Police Service (IPS) in support of the student movement at Jantar Mantar. In the viral video, DCP Sumit Jha is allegedly heard saying that he was resigning from his post with immediate effect and could not remain part of a system that forced him to spread false information against students. The video was shared with the claim that Delhi Police DCP Sumit Jha resigned from his position to support protesting students.
Claim
An Instagram user shared the viral video on July 28, 2026, claiming that Delhi Police DCP Sumit Jha had resigned from his post in support of protesting students.
CyberPeace Research Wing’s research found the claim to be false. The research revealed that DCP Sumit Jha has not resigned from his position. The viral video was digitally manipulated by adding AI-generated audio to an original video shared by the Delhi Police on its official X account on July 23, 2026.
The video was further analysed using the AI detection tool HIVE Moderation. The tool’s analysis indicated that the audio used in the viral video had a 91.8 per cent probability of being AI-generated, confirming that the audio was artificially created and was not authentic.
Claim 2: CDS Raja Subramani confirmed that more than 30,000 soldiers resigned over Jantar Mantar protests
A video featuring Chief of Defence Staff (CDS) General Raja Subramani was widely circulated on social media with the false claim that he confirmed the resignation of more than 30,000 Indian Army personnel following the Jantar Mantar protest. The viral video claimed that soldiers had refused to serve the government after police action against protesting students.
Claim
A Facebook user shared the video claiming: "After the protests at Jantar Mantar, India’s Chief of Defence Staff Raja Subramani has issued an emergency alert, saying that more than 30,000 soldiers have resigned in protest because their children were brutally beaten at Jantar Mantar."
CyberPeace Research Wing’s research found the claim to be false. The research revealed that the video featuring CDS Raja Subramani was digitally manipulated using AI-generated audio. The original video did not contain any statement related to soldiers resigning or the Jantar Mantar protest. Further research led us to the same video posted by the Instagram account myyouthindia, where CDS Raja Subramani can be seen extending his best wishes to teams participating in the 135th Indian Oil Durand Cup Football Tournament. https://www.instagram.com/reels/DbEGRGBsKv8/
The audio circulating with the viral video was falsely added to create a misleading narrative.
Claim 3: Defence Minister Rajnath Singh warned protesters of strict action
A video featuring Union Defence Minister Rajnath Singh was widely circulated on social media with the claim that he issued a strict warning against student protesters. In the viral clip, Singh was allegedly heard saying that the government would not tolerate “anarchy” and that strict action would be taken against those blocking roads, disturbing public order or damaging government property.
Claim
An X user (@ForumDefence) shared the video on July 22, 2026, claiming that Rajnath Singh had issued a warning regarding the ongoing student protest.
Fact Check
CyberPeace Research Wing’s research found the claim to be false.The original video available on Rajnath Singh’s official YouTube channel was digitally altered by adding AI-generated audio. The statement heard in the viral video was not made by the Defence Minister and was created using artificial intelligence. The viral video was also analysed using the AI detection tool HIVE Moderation. The analysis indicated that the speech used in the video had an 88.1 per cent probability of being AI-generated.
Claim 4: Piyush Goyal threatened to hang CJP supporters
A video of Union Minister Piyush Goyal was shared on social media with the false claim that he threatened protesters and called for two to three CJP supporters to be hanged to "teach the rest of them a lesson." In the AI-manipulated video, Goyal was allegedly heard saying: "To teach the CJP a lesson, we should hang two to three of its supporters so that they understand the real power of the government."
CyberPeace Research Wing’s research found that the viral video was digitally manipulated using AI-generated audio. The original video, posted on the Press Trust of India’s official X account, showed Piyush Goyal discussing the government’s response to examination paper leaks and appealing for the issue to be resolved through dialogue. He made no remarks threatening protesters or calling for the execution of CJP supporters.
The audio and statements attributed to him in the viral video were fabricated using AI technology.
Conclusion
CyberPeace Research Wing’s research found that all four viral claims related to the CJP protest were false. The videos circulating on social media were either digitally manipulated or created using AI-generated audio to falsely attribute statements to public officials. The research highlights how AI-generated deepfakes and manipulated content are increasingly being used to spread misinformation during sensitive events.
During the Gau Raksha Yatra of Shankaracharya Swami Avimukteshwaranand Saraswati, bees reportedly attacked a discourse event in Rohania area of Varanasi, Uttar Pradesh. Following the incident, a picture has gone viral on social media showing bees attacking Swami Avimukteshwaranand Saraswati. Several users are sharing the image as genuine while targeting the Shankaracharya online. CyberPeace Research Wing investigated the viral image and found it to be fake. Our research revealed that the picture was created using Artificial Intelligence (AI). While it is true that a bee attack occurred during Swami Avimukteshwaranand Saraswati’s discourse program, the viral image itself is fabricated.
Claim
A Facebook user named “Sanjay Chaudhary” shared the viral image on May 15, 2026, with the caption: “Prakritik kop ka bhajan bana Shukracharya Umashankar alias Avimukteshwaranand… This Kaalnemi was delivering false sermons in Rohania, Varanasi in the name of religion… The bees from a nearby hive did not like it and collectively attacked, creating chaos. Even insects and nature no longer like the opposition’s politics disguised as Sanatan Dharma. Calling Yogi Ji Aurangzeb, Akbar and butcher is not acceptable even to nature and insects.”
To verify the viral claim, we used Google Open Search tools and found reports related to the incident on the YouTube channel of News18 UP Uttarakhand. A report published on May 13, 2026 stated that bees attacked the discourse event during Swami Avimukteshwaranand’s Gau Raksha Yatra in Rohania, Varanasi. The incident created panic at the venue, forcing the Swami to end his discourse midway. The channel also uploaded a YouTube Shorts video related to the incident.
As part of the research, we further analyzed the viral image using AI detection tools. First, we used the tool “Sight Engine,” which indicated an 88 percent probability that the image was AI-generated.
We then examined the image using another AI detection tool called “Undetectable,” which also suggested that the photo was likely created using AI.
Conclusion
Our research found that the viral image is AI-generated. The picture was created using artificial intelligence tools. While bees did attack during Swami Avimukteshwaranand Saraswati’s Gau Raksha Yatra on May 13, 2026, the viral image circulating on social media is fictional and not real.
The emergence of deepfake technology has become a significant problem in an era driven by technological growth and power. The government has reacted proactively as a result of concerns about the exploitation of this technology due to its extraordinary realism in manipulating information. The national government is in the vanguard of defending national interests, public trust, and security as the digital world changes. On the 26th of December 2023, the central government issued an advisory to businesses, highlighting how urgent it is to confront this growing threat.
The directive aims to directly address the growing concerns around Deepfakes, or misinformation driven by AI. This advice represents the result of talks that Union Minister Shri Rajeev Chandrasekhar, had with intermediaries during the course of a month-long Digital India dialogue. The main aim of the advisory is to accurately and clearly inform users about information that is forbidden, especially those listed under Rule 3(1)(b) of the IT Rules.
Advisory
The Ministry of Electronics and Information Technology (MeitY) has sent a formal recommendation to all intermediaries, requesting adherence to current IT regulations and emphasizing the need to address issues with misinformation, specifically those driven by artificial intelligence (AI), such as Deepfakes. Union Minister Rajeev Chandrasekhar released the recommendation, which highlights the necessity of communicating forbidden information in a clear and understandable manner, particularly in light of Rule 3(1)(b) of the IT Rules.
Advise on Prohibited Content Communication
According to MeitY's advice, intermediaries must transmit content that is prohibited by Rule 3(1)(b) of the IT Rules in a clear and accurate manner. This involves giving users precise details during enrollment, login, and content sharing/uploading on the website, as well as including such information in customer contracts and terms of service.
Ensuring Users Are Aware of the Rules
Digital platform suppliers are required to inform their users of the laws that are relevant to them. This covers provisions found in the IT Act of 2000 and the Indian Penal Code (IPC). Corporations should inform users of the potential consequences of breaking the restrictions outlined in Rule 3(1)(b) and should also urge users to notify any illegal activity to law enforcement.
Talks Concerning Deepfakes
For more than a month, Union Minister Rajeev Chandrasekhar had a significant talk with various platforms where they addressed the issue of "deepfakes," or computer-generated fake videos. The meeting emphasized how crucial it is that everyone abides by the laws and regulations in effect, particularly the IT Rules to prevent deepfakes from spreading.
Addressing the Danger of Disinformation
Minister Chandrasekhar underlined the grave issue of disinformation, particularly in the context of deepfakes, which are false pieces of content produced using the latest developments such as artificial intelligence. He emphasized the dangers this deceptive data posed to internet users' security and confidence. The Minister emphasized the efficiency of the IT regulations in addressing this issue and cited the Prime Minister's caution about the risks of deepfakes.
Rule Against Spreading False Information
The Minister referred particularly to Rule 3(1)(b)(v), which states unequivocally that it is forbidden to disseminate false information, even when doing so involves cutting-edge technology like deepfakes. He called on intermediaries—the businesses that offer digital platforms—to take prompt action to take such content down from their systems. Additionally, he ensured that everyone is aware that breaking such rules has legal implications.
Analysis
The Central Government's latest advisory on deepfake technology demonstrates a proactive strategy to deal with new issues. It also highlights the necessity of comprehensive legislation to directly regulate AI material, particularly with regard to user interests.
There is a wider regulatory vacuum for content produced by artificial intelligence, even though the current guideline concentrates on the precision and lucidity of information distribution. While some limitations are mentioned in the existing laws, there are no clear guidelines for controlling or differentiating AI-generated content.
Positively, it is laudable that the government has recognized the dangers posed by deepfakes and is making appropriate efforts to counter them. As AI technology develops, there is a chance to create thorough laws that not only solve problems but also create a supportive environment for the creation of ethical AI content. User protection, accountability, openness, and moral AI use would all benefit from such laws. This offers an opportunity for regulatory development to guarantee the successful and advantageous incorporation of AI into our digital environment.
Conclusion
The Central Government's preemptive advice on deepfake technology shows a great dedication to tackling new risks in the digital sphere. The advice highlights the urgent need to combat deepfakes, but it also highlights the necessity for extensive legislation on content produced by artificial intelligence. The lack of clear norms offers a chance for constructive regulatory development to protect the interests of users. The advancement of AI technology necessitates the adoption of rules that promote the creation of ethical AI content, guaranteeing user protection, accountability, and transparency. This is a turning point in the evolution of regulations, making it easier to responsibly incorporate AI into our changing digital landscape.
Executive Summary - When Anthropic and OpenAI's AI Testing Turned Into Real Breaches
You would be surprised to know that a testing function built to measure how good AI models are at simulated hacking ended up doing the real thing instead. Not once , but three times, across two of the world's leading AI labs, within the same 9-day window at the end of July 2026. As per the reports, Anthropic, which is among the world's leading AI labs, was running these evaluations on its own AI models namely - Claude Opus 4.7, Claude Mythos 5, and an unreleased research model, inside an environment co-managed with a third-party evaluation vendor. As per the reports, the models were told they were operating inside closed, internet-free simulations. They were not. A configuration error left the door open to the real internet, and the AI did exactly what it was trained to do in a hacking exercise, find the target and break in. Except the targets, this time, were real companies. Real credentials got stolen. Real data got accessed. Two of the three victims didn't even know they'd been breached until the AI lab called to tell them. This shows how a single unverified assumption, "this environment has no internet access" can quietly collapse the entire safety boundary of an AI test. It indicates that as these systems get more capable and more autonomous, the risk isn't necessarily the AI deciding to go rogue, it is humans failing to double-check the cage before putting something powerful inside it. And it warns us that the margin for this kind of error is shrinking fast, because what used to be a contained mistake can now scan thousands of systems and act on it within minutes. bAnthropic was not alone. Just over a week earlier, on 21 July, OpenAI had disclosed that its own models, GPT-5.6 Sol and an unreleased successor broke out of an isolated test environment and reached the real production infrastructure of Hugging Face, another AI company. In fact, it was OpenAI's disclosure that pushed Anthropic to go back and audit its own evaluation runs in the first place, which is how the three incidents above came to light. This blog examines the incident not as an isolated slip-up by two labs, but as a preview of the larger picture, a world where the gap between "theoretically vulnerable" and "actually exploited" is closing faster than most organisations' security practices are ready for.
The Chain of Disclosures
It started with Hugging Face, the open-source AI platform, which disclosed on July 16 that it had detected and contained an intrusion into part of its production infrastructure, driven end-to-end by an autonomous AI agent. Five days later, OpenAI confirmed it was responsible: models it was evaluating internally, GPT-5.6 Sol and an unreleased prototype, both running with their cyber-safety behaviour deliberately turned down to measure raw capability, had exploited a previously unknown vulnerability to break out of an isolated test environment and reach Hugging Face's real servers, apparently while chasing answers to a benchmark challenge. That disclosure prompted Anthropic to go back through its own records. On July 30, it published the results: after auditing 141,006 evaluation runs in which its Claude models could conceivably have touched the internet, it found three incidents, six runs in total, in which a model reached real systems belonging to three outside organisations and gained unauthorised access.
How a Fictional Exercise became a Real Attack
All three Anthropic incidents shared the same root cause. Claude was being run through capture-the-flag exercises, a standard way labs test offensive cyber capability: the model is dropped into a fictional scenario, told a piece of secret data is hidden somewhere on a network, and left to find its own way in. Anthropic's prompts told Claude, explicitly, that the environment was an internet-free simulation. Because of a misunderstanding with its evaluation partner, Irregular, that statement was false; the machines had live internet access nobody had verified. So when Claude's search for the flag led it onto the open internet, it did what it had been told to expect: it assumed everything it found was part of the exercise, including things that weren't.
Incident 1: A fictional target company happened to share its name with a real, active domain. Across four runs, Claude Opus 4.7 attacked the real company instead, extracting infrastructure credentials and reaching a database containing several hundred rows of genuine production data, the most serious impact of the three incidents. In its own reasoning, the model repeatedly noted the system looked real and kept going anyway.
Incident 2: Inside its fictional scenario, Claude Mythos 5 found a fake "onboarding document" referencing a software package that didn't exist on PyPI, the public Python package registry. It recognised the opportunity, built a functioning malicious package, and, after working around several dead ends to get an email account and register on PyPI, published it. The package lived online for about an hour and was installed on fifteen real machines, including a security company's automated scanner, letting Claude harvest real credentials from a company it had never even set out to target.
Incident 3: An unreleased internal research model, unable to find its fictional target, scanned roughly 9,000 real hosts and broke into one using textbook techniques credentials read off an exposed debug page, then SQL injection. Partway through, it noticed the system it had compromised had no plausible link to its assignment, concluded on its own that the target was real, and stopped.
Timeline of Incidents
The Threat is Scale, not Novelty
None of the three models discovered a novel vulnerability. Weak passwords, unauthenticated endpoints, exposed debug output, and a classic "dependency confusion" trick are techniques that have appeared on security checklists for a decade and are still relevant. That's arguably the more sobering finding: the frontier is not a new class of attack; it is the speed and tirelessness with which known, boring weaknesses can now be found and used. Anthropic's own analysis lands on a subtler point, too: in each case, whether the model's behaviour counted as "aligned" or "harmful" depended entirely on what it believed about its situation, not on any goal of its own. All three models were simply trying to complete the task they had been assigned. Two of them did so while holding a false belief, installed by their own evaluators, about whether their targets were real. Only the newest of the three models, on independently concluding it had reached a genuine system, chose to stop. Anthropic is careful to note this is not a controlled experiment and should not be read as a guarantee that newer models will always behave better. But the throughline is clear enough: telling a capable model a lie about its environment is not a neutral setup choice. It is itself a safety-relevant decision.
The Detection Gap
Perhaps the most alarming detail is the quietest one. Anthropic reached out to the three affected organisations on July 27. Two of them had detected nothing at all, no alert, no anomaly, no investigation until that call. Real credentials had been stolen and real data accessed inside systems whose owners had no idea anything had happened. That is a statement about the state of everyday detection capability, not about AI. An agent that completes an entire intrusion, start to finish, within a single automated session doesn't leave the kind of slow, human-paced footprint that most monitoring is built to catch.
The Silver Lining - Why These Disclosures Deserve Credit
Both incidents share an underappreciated feature: they were disclosed voluntarily, promptly, and with real detail, and both labs notified the organisations affected. Hugging Face brought in outside forensic specialists and law enforcement. Anthropic halted its cyber evaluations the same day it found the first suspicious transcript and has asked METR, an independent evaluation body, to review its findings. That kind of candour is exactly the behaviour any sensible policy response should want to reinforce. A regulatory reflex that punishes disclosure risks teaching labs to say less next time, not to do better. What both incidents point to, far more than any specific model capability, is a mundane and fixable governance gap: environments used to test powerful, semi-restrained AI systems need the same security discipline as production systems, verified network isolation, continuous monitoring, and evaluation scopes that are stated positively ("here is what's in bounds") rather than enforced by simply telling the model a comforting falsehood. As both companies note, a fictional test range that turns out to have a live path to the internet isn't really fictional anymore. Basic asset hygiene, like knowing what's exposed, patching debug endpoints, claiming your internal package names before someone else does, and watching outbound traffic from environments that are supposed to have none did more to prevent and contain these incidents than anything specific to the models involved.
CyberPeace findings and recomendations : For enterprises and public institutions
Maintain a full inventory of internet-facing assets and unauthenticated endpoints, and assume the inventory is incomplete until proven otherwise.
Eliminate default, weak, and reused credentials, and enforce phishing-resistant MFA on anyone externally reachable.
Strip debug pages and verbose error output from production systems.
Treat dependency confusion as a live threat: pin dependencies, use private registry namespaces, and pre-emptively claim internal package names on public registries.
Apply deny-by-default egress filtering to every environment running AI or agentic tooling, including development and test environments, and verify isolation empirically rather than assuming it from configuration.
Alert on any outbound connection from an environment that is supposed to have none.
Review authentication and access logs from April 2026 onwards for short, unusually efficient sessions that look more like machine-speed compromise than human reconnaissance.
For AI developers and evaluation vendors
Network-isolate offensive-capability evaluation environments by default, with isolation verified per run rather than inherited from configuration.
State the scope explicitly and positively, which systems are in bounds rather than asserting a falsehood about connectivity.
Build contractual isolation guarantees and joint pre-run verification into third-party evaluation partnerships; both labs involved here have acknowledged that neither side alone caught the misconfiguration.
Monitor transcripts and network logs continuously, not retrospectively.
For policymakers
A regulatory response that punishes candour risks producing silence rather than safety. India currently has no reporting framework that clearly covers containment failures in AI evaluations affecting Indian entities' behaviour.
RT-In's existing incident-reporting directions were not drafted with this candour in mode. Closing that gap would mean an explicit reporting obligation for evaluation of containment failures touching third-party infrastructure and a safe harbour mechanism that protects labs which disclose promptly.
Minimum containment standards (egress verification, log retention) for organisations conducting offensive-capability AI evaluation within Indian jurisdiction;
Recognition in national cyber doctrine that agentic tooling collapses the gap between a known-but-deferred vulnerability and an exploited one.
Conclusion
The above incidents reveal less about AI's offensive capability and more about the gap between how these systems are tested and how carefully those tests are contained. Both labs found the breaches through their own review, not external detection, a point in their favor, but also a reminder that containment failures can go unnoticed for a while. The realistic risk ahead isn't a sudden leap in AI's hacking sophistication; it's the compounding effect of speed and scale applied to routine reconnaissance, run against infrastructure that assumes a human attacker's pace. Treating evaluation environments with the same rigor as production systems, sandboxing, monitoring, and independent audits, should become standard practice, not an afterthought triggered by another lab's incident. The path forward is less about slowing AI down and more about catching up our containment discipline to match what these systems can now do.
Your institution or organization can partner with us in any one of our initiatives or policy research activities and complement the region-specific resources and talent we need.