AI Security Trend Roundup: Oct 02, 2026
Items from the week ending Oct 02, 2026. arXiv items announced Oct 2, 2026 (IDs: 2610.xxxxx). 23 items from 2 sources.
This digest credits every source by name and links directly to each original post. Items are selected automatically by keyword and feed filters. All rights and attribution belong to the original authors.
Academic & Research
Source: arXiv cs.CR, Oct 2026 arXiv:2610.00125v1 Announce Type: new Abstract: One-Pixel Attacks (OPAs) represent one of the most extreme demonstrations of adversarial fragility in deep learning, where modifying a single pixel can reliably induce high-confidence misclassification across domains such as medical
Source: arXiv cs.CR, Oct 2026 arXiv:2610.00126v1 Announce Type: new Abstract: Agent developers increasingly compare prompts, tools, policies, and diagnosis algorithms through simulator-grounded verifiers. A verifier can nevertheless make a solver comparison vacuous: if its probes or predicates encode the targ
Source: arXiv cs.CR, Oct 2026 arXiv:2610.00136v1 Announce Type: new Abstract: This paper presents a reproducible, educational study of evasion attacks in image classification and text classification. A compact convolutional network trained on MNIST reached 98.63% clean test accuracy and was evaluated under tw
Source: arXiv cs.CR, Oct 2026 arXiv:2610.00151v1 Announce Type: new Abstract: Agent deployments increasingly combine language-model inference with retrieval, delegation, tool execution, external-system access, and human approval. Security-relevant deviations can therefore emerge across an evolving process rat
Source: arXiv cs.CR, Oct 2026 arXiv:2610.00327v1 Announce Type: new Abstract: Tool-using agents can expose citations and execution logs while leaving a critical association unaudited: whether the claim shown to a user is the claim emitted by the committed execution and supported by the cited source. A valid c
Source: arXiv cs.CR, Oct 2026 arXiv:2610.00347v1 Announce Type: new Abstract: Self-modifying AI agents can replace, fork, and roll back identity-bearing software while descendants remain executable. Per-successor authorization does not constrain the resulting population: siblings may duplicate quotas, combine
Source: arXiv cs.CR, Oct 2026 arXiv:2610.00354v1 Announce Type: new Abstract: AI agents that control wallets read attacker-reachable content, so they can be steered into proposing harmful transactions. The usual last line of defense is a pre-signing check: a static allowlist, an LLM reviewer, or a transaction
Source: arXiv cs.CR, Oct 2026 arXiv:2610.00382v1 Announce Type: new Abstract: Model quantization reduces the numerical precision of neural network weights and activations to lower storage and computational costs. Model inversion attacks recover or reconstruct sensitive training data or inference inputs from m
Source: arXiv cs.CR, Oct 2026 arXiv:2610.00392v1 Announce Type: new Abstract: Agent interaction protocols such as ACP and A2A have moved LLM-based agents toward multi-agent collaboration, introducing new security threats. A task sent by a remote peer over A2A is treated as a legitimate request, providing a na
Source: arXiv cs.CR, Oct 2026 arXiv:2610.00787v1 Announce Type: new Abstract: A correctly governed LLM agent can reach a state in which neither continuing execution nor automatically halting is admissible: the system has detected a persistent failure of its observability or drift-detection layer, but cannot i
Source: arXiv cs.CR, Oct 2026 arXiv:2610.00839v1 Announce Type: new Abstract: Large language models (LLMs) deployed through text-only APIs face model extraction risks, as adversaries can collect their responses to train surrogates that reproduce their capabilities. While prior work has developed diverse attac
AI News & Commentary
Source: Simon Willison, quoting Matthew Green, Oct 01 [...] Put these pieces together and you have the two halves of a worm: a payload that hijacks the agent, and an agent that will carry the payload to the next agent. Agents in separately-isolated sandboxes discovered that they could leave instructions for each other in a shared pa
Source: Simon Willison, Sep 30 I visited the Museum of the City of New York today and got to see He Built This City: Joe Macken’s Model, the 50 by 27 feet model of the city built over a 21 year period from balsa wood and cardboard. It exceeded my already high expectations. The exhibition closes on 12th October s
Source: Simon Willison, Sep 29 Anthropic Frontier Red Team, quoted by Simon Willison: We evaluate several models on 100 tasks from the [internal Binary Exploitation benchmark] (selected at random), and find that GLM-5.3 develops full control flow hijacks in 4% of the trials; Claude Mythos Preview did so in 6%. Although GLM-5.3 performs below Claude Mythos Preview
Source: Simon Willison, Sep 29 My comment on GPT 6.1 Sol: Near-Astra intelligence for a fifth of the price, Hacker News. I'm a bit late with the pelicans because I was live-blogging the keynote: https://simonwillison.net/2026/Sep/29/openai-devday-2026-liv... Here they are for GPT-6.1-Sol: https://tools.s
Source: Simon Willison, Sep 29 Tool: Photo Scrubber, local face blur & metadata removal. I took a photograph of some protesters, then thought about how I don't like sharing photographs of strangers with identifiable faces. I had GPT-6 Astra build this experimental tool that would identify faces and automatical
Source: Simon Willison, Sep 29 I'm at OpenAI DevDay today, in Fort Mason, San Francisco. Same as last year I'll be live blogging the keynote and some other notes during the day. OpenAI gave me a free ticket and a seat in the "creator" area for the keynote. Tags: ai, openai, generative-ai, llms, coding-agents,
Source: Simon Willison, Sep 28 Claude Sonnet 5.5 New Sonnet model from Anthropic today. Anthropic says it "runs 30%+ faster, and costs up to 30% less for most work" - it's priced the same as Sonnet 5 but appears to beat it on every benchmark, and should be cheaper to run as well. Here are some pelicans riding bicycl
Source: Simon Willison, Sep 28 To say that we were surprised at the jump and suddenness of the capabilities of our models when it came to “cyber” or “swarming” or “message boards” or anything else related to the incidents is an understatement. Security posture takes time to develop. It’s not just about hardeni
Source: Simon Willison, Sep 27 On Friday I gave the closing keynote at the WeAreDevelopers World Congress North America in San Jose. I tied together the key trends from the past year into a chronological exploration of everything that happened in 2026. The video is on YouTube; here are my annotated slides and
Source: Simon Willison, Sep 27 My comment on S3 Is the Future, S3 Is the Past, Hacker News. One thing I find notable about S3 today is that, while it used to drop in price reasonably often, there hasn't been a price drop in a full decade: 2006-03-14 $0.150/GB-month 2010-11-01 $0.140/GB-month 2012-02-01 $
Source: Simon Willison, Sep 27 Tool: Bluesky reply bot checker Automated reply bots on Twitter are a scourge - as someone with a decent number of followers I attract a swarm of these, such that anything I post there attracts dozens of mindless automated replies. They've started manifesting on Bluesky as well.
Source: Simon Willison, Sep 26 Tool: Kākāpō Party I presented a closing keynote for the WeAreDevelopers World Congress North America yesterday. As a STAR moment I decided to weave in references to the record breaking kākāpō breeding season we had in 2026. For my closing slide I wanted to celebrate, and I had s
Source List
Sources in this roundup, credited to their original authors/organizations:
- Simon Willison, feed:
https://simonwillison.net/atom/everything/ - arXiv cs.CR, feed:
http://export.arxiv.org/rss/cs.CR
Explore More
FixTheVuln Store
Studying for Security+? Get the Study Planner
Structured study planners for CompTIA certifications. Domain trackers, time blocking, and exam strategies.
Shop Security+ PlannerAlso available: CompTIA A+, Network+, CySA+, PenTest+
CyberFolio
Building cybersecurity skills? Track them in one place.
Build a shareable cybersecurity portfolio that highlights your certifications, projects, and skills, free.
Build Your Portfolio →