AI Newsway

GitHub's Open-Source AI Audit Agent Found 24 Android Bugs β€” and Kept Misjudging How Bad They Were

A silent location leak in OsmAnd and a one-tap Wikipedia account takeover came out of taskflows any developer can run on their own repo

|5 min read0
AI Summary
GitHub Security Lab has reported 24 Android vulnerabilities found with custom taskflows running on its open-source Taskflow Agent. Disclosed examples include an exported activity in OsmAnd that let any app silently rewrite map settings and leak a user's routes, and a Wikipedia deeplink flaw enabling account takeover from one tap. Researcher Kevin Stubbings said the model found bugs well but repeatedly misjudged severity, so human review remains required.
Source code on a developer's screen, the material GitHub Security Lab's open-source taskflow agent audits to surface Android vulnerabilities
Source code on a developer's screen, the material GitHub Security Lab's open-source taskflow agent audits to surface Android vulnerabilities

GitHub's Security Lab has reported 24 vulnerabilities in Android applications using an open-source AI agent that any developer can point at their own repository. Researcher Kevin Stubbings described the work in a post on the GitHub Blog, walking through two already-disclosed findings: a silent location leak in the navigation app OsmAnd, and a one-tap account takeover in the Wikipedia Android app.

Key takeaways

  • Twenty-four Android vulnerabilities have been found and reported with custom taskflows built on GitHub Security Lab's open-source Taskflow Agent.
  • An exported activity in OsmAnd, an app with more than 10 million Play Store downloads, let any installed app rewrite map settings and stream a user's tile coordinates and routes to an attacker's server.
  • The model proved much better at finding bugs than at rating them, producing persistent low-severity noise and false positives that only human review caught.

How the taskflows were tuned for Android

The agent itself is generic; the value sat in the prompts. Stubbings added a taskflow, gather_mobile_entry_point_info.yaml, that splits entry points into mobile and non-mobile buckets, so a repository holding a phone app, a web server and a desktop client does not leave the model reasoning about the wrong attack surface.

He then edited classify_application_local.yaml to name vulnerability classes outright β€” confused deputy issues, insecure broadcasts β€” rather than trusting recall. Because model output is non-deterministic, the strict prompt ran repeatedly alongside a broader one: the narrow version catches the obvious, the broad version finds the unexpected.

What the OsmAnd flaw allowed

Three of the 24 reports sit in OsmAnd, and the most serious reads as an architecture bug rather than a coding slip.

  1. MapActivity, the screen handling deeplinks and settings imports, is exported β€” reachable by any other app on the device.
  2. Its import path reads intent extras such as silent_import, replace and export_type_list_key, which the developers expected only over an in-process AIDL channel. Android offers no way to restrict which extras an external caller attaches.
  3. A permissionless app can therefore write settings with no notification and no confirmation β€” including the template OsmAnd uses to build map tile URLs.
  4. Pointed at an attacker's host, that template leaks the coordinates of every tile loaded and every route planned, while the server proxies genuine OpenStreetMap imagery so the map looks untouched.

No memory-safety error or unusual device configuration is involved. This is the class of logic flaw static analysers rarely flag and reviewers rarely have time to trace end to end.

A deeplink check that trusted the wrong string

The Wikipedia finding ends in an account takeover from one tapped link, and the cause is a single string comparison: its deeplink handler validated the URL authority with endsWith against the base domain instead of matching it. A wikipedia:// URL aimed at a lookalike such as evil-wikipedia.org therefore passed and rendered inside the app's WebView, running attacker JavaScript in a context the app treats as its own.

A second check in the cookie manager repeated the mistake, letting that page read session cookies valid across every Wikimedia property.

Where the model fell short

Stubbings was blunt about the weak spot: severity estimation. The model kept reporting low-impact issues even when told not to, and misjudged real-world impact when a mitigating factor cancelled the exploit β€” a path traversal into external storage is worthless if the app prioritises internal storage for the same data. Forcing it to build a working proof of concept surfaces some of those cases, at the cost of extra runs on bugs that may not matter.

The compensating strength was API knowledge: the model reliably separated security-relevant function behaviour, such as Go's path.Clean versus filepath.Clean, and most generated proofs of concept needed little editing. As Help Net Security summarised it, every finding still needs a reviewer who knows mobile apps.

Running it on your own project

The taskflows live in the seclab-taskflows repository, set up to launch in a Codespace; ./scripts/audit/run_mobile.sh myorg/myrepo starts an audit that takes an hour or two on a medium codebase and writes flagged rows to a SQLite has_vulnerability column. It needs a Copilot licence and consumes premium model requests, and GitHub warns the token bill is not trivial. It is the constructive counterpart to a trend we have covered from the other side: attackers running agents that rewrite malware until scanners stop catching it.

FAQ

Is GitHub's AI security agent free to use?

The Taskflow Agent and the example taskflows are open source, but running them is not free in practice. A GitHub Copilot licence is required and a single audit issues a large number of premium model requests, which GitHub explicitly flags as a meaningful cost on larger repositories.

Have the OsmAnd and Wikipedia vulnerabilities been disclosed?

Yes β€” both were reported through the normal process and publicly described only after disclosure. GitHub Security Lab publishes the rest of the 24 findings on its advisories page as each one is released.

Can an AI agent replace a human security reviewer?

Not on this evidence. The agent generated useful leads and working proofs of concept, but it over-reported low-severity issues and mis-rated impact where mitigating factors applied, so a researcher familiar with the platform has to validate each finding.

How do you feel about this article?

SJ

Discussion

Sign in to post
Loading...

Related articles

Vercel Disabled AVIF Platform-Wide. The Bug Was Three Layers Below Next.js
Developer Tools

Vercel Disabled AVIF Platform-Wide. The Bug Was Three Layers Below Next.js

Vercel traced a reported Next.js RCE to libheif, disabled AVIF platform-wide on August 13, and coordinated fixes across sharp, libvips and libheif by August 25.

Seung Jung10 days ago
ZCode Packaged 42,411 Files Per Snapshot. Only Z.ai Could Decrypt Them.
Developer Tools

ZCode Packaged 42,411 Files Per Snapshot. Only Z.ai Could Decrypt Them.

A reverse-engineering report found Z.ai's ZCode app shipping full Git histories to Alibaba Cloud under encryption only the company's servers could unwrap.

Seung Jung10 days ago
Meta Open-Sources Astryx, a React Design System Agents Can Query
Developer Tools

Meta Open-Sources Astryx, a React Design System Agents Can Query

Meta released Astryx in June, a React design system that matured for eight years inside the company's internal monorepo, as a public beta under the MIT license....

Seung Jung15 days ago
Researchers Reached OpenAI's Internal Repo Through a Forum Image Bug
Developer Tools

Researchers Reached OpenAI's Internal Repo Through a Forum Image Bug

Hacktron AI chained a libheif heap overflow with an OpenAI SSO flaw to reach employee Codex accounts and the openai/openai monorepo. Both bugs are patched.

Seung Jung11 days ago
832,378 Lines of Rust in 14.5 Weeks: Inside an Agent-Run Rewrite
Developer Tools

832,378 Lines of Rust in 14.5 Weeks: Inside an Agent-Run Rewrite

GitHub converted 430,000 lines of TypeScript into 832,378 lines of Rust in 14.5 weeks using coding agents. Memory use fell from 1,383MB to 126MB.

Seung Jung10 days ago
Microsoft's Record 966-Flaw Patch Month Moves the Bottleneck to Defenders
Developer Tools

Microsoft's Record 966-Flaw Patch Month Moves the Bottleneck to Defenders

Microsoft fixed 966 vulnerabilities in September, pushing its 2026 total near 2,750. Security teams say triage, not discovery, is now the hard part.

Seung Jung13 days ago