GitHub's Security Lab has reported 24 vulnerabilities in Android applications using an open-source AI agent that any developer can point at their own repository. Researcher Kevin Stubbings described the work in a post on the GitHub Blog, walking through two already-disclosed findings: a silent location leak in the navigation app OsmAnd, and a one-tap account takeover in the Wikipedia Android app.
Key takeaways
- Twenty-four Android vulnerabilities have been found and reported with custom taskflows built on GitHub Security Lab's open-source Taskflow Agent.
- An exported activity in OsmAnd, an app with more than 10 million Play Store downloads, let any installed app rewrite map settings and stream a user's tile coordinates and routes to an attacker's server.
- The model proved much better at finding bugs than at rating them, producing persistent low-severity noise and false positives that only human review caught.
How the taskflows were tuned for Android
The agent itself is generic; the value sat in the prompts. Stubbings added a taskflow, gather_mobile_entry_point_info.yaml, that splits entry points into mobile and non-mobile buckets, so a repository holding a phone app, a web server and a desktop client does not leave the model reasoning about the wrong attack surface.
He then edited classify_application_local.yaml to name vulnerability classes outright β confused deputy issues, insecure broadcasts β rather than trusting recall. Because model output is non-deterministic, the strict prompt ran repeatedly alongside a broader one: the narrow version catches the obvious, the broad version finds the unexpected.
What the OsmAnd flaw allowed
Three of the 24 reports sit in OsmAnd, and the most serious reads as an architecture bug rather than a coding slip.
MapActivity, the screen handling deeplinks and settings imports, is exported β reachable by any other app on the device.- Its import path reads intent extras such as
silent_import,replaceandexport_type_list_key, which the developers expected only over an in-process AIDL channel. Android offers no way to restrict which extras an external caller attaches. - A permissionless app can therefore write settings with no notification and no confirmation β including the template OsmAnd uses to build map tile URLs.
- Pointed at an attacker's host, that template leaks the coordinates of every tile loaded and every route planned, while the server proxies genuine OpenStreetMap imagery so the map looks untouched.
No memory-safety error or unusual device configuration is involved. This is the class of logic flaw static analysers rarely flag and reviewers rarely have time to trace end to end.
A deeplink check that trusted the wrong string
The Wikipedia finding ends in an account takeover from one tapped link, and the cause is a single string comparison: its deeplink handler validated the URL authority with endsWith against the base domain instead of matching it. A wikipedia:// URL aimed at a lookalike such as evil-wikipedia.org therefore passed and rendered inside the app's WebView, running attacker JavaScript in a context the app treats as its own.
A second check in the cookie manager repeated the mistake, letting that page read session cookies valid across every Wikimedia property.
Where the model fell short
Stubbings was blunt about the weak spot: severity estimation. The model kept reporting low-impact issues even when told not to, and misjudged real-world impact when a mitigating factor cancelled the exploit β a path traversal into external storage is worthless if the app prioritises internal storage for the same data. Forcing it to build a working proof of concept surfaces some of those cases, at the cost of extra runs on bugs that may not matter.
The compensating strength was API knowledge: the model reliably separated security-relevant function behaviour, such as Go's path.Clean versus filepath.Clean, and most generated proofs of concept needed little editing. As Help Net Security summarised it, every finding still needs a reviewer who knows mobile apps.
Running it on your own project
The taskflows live in the seclab-taskflows repository, set up to launch in a Codespace; ./scripts/audit/run_mobile.sh myorg/myrepo starts an audit that takes an hour or two on a medium codebase and writes flagged rows to a SQLite has_vulnerability column. It needs a Copilot licence and consumes premium model requests, and GitHub warns the token bill is not trivial. It is the constructive counterpart to a trend we have covered from the other side: attackers running agents that rewrite malware until scanners stop catching it.
FAQ
Is GitHub's AI security agent free to use?
The Taskflow Agent and the example taskflows are open source, but running them is not free in practice. A GitHub Copilot licence is required and a single audit issues a large number of premium model requests, which GitHub explicitly flags as a meaningful cost on larger repositories.
Have the OsmAnd and Wikipedia vulnerabilities been disclosed?
Yes β both were reported through the normal process and publicly described only after disclosure. GitHub Security Lab publishes the rest of the 24 findings on its advisories page as each one is released.
Can an AI agent replace a human security reviewer?
Not on this evidence. The agent generated useful leads and working proofs of concept, but it over-reported low-severity issues and mis-rated impact where mitigating factors applied, so a researcher familiar with the platform has to validate each finding.






