Here's your daily roundup of the most relevant AI and ML news for September 02, 2026. Today's digest includes 1 security-focused story. We're also covering 5 research developments. Click through to read the full articles from our curated sources.
Security & Safety
1. Researchers Use Claude to Port Pre-Auth RCE Exploit From One PLC Model to Another
Forescout Research - Vedere Labs said it used Anthropic's Claude to port a working pre-authentication remote code execution (RCE) exploit from one WAGO programmable logic controller (PLC) to another, executing attacker-supplied ARM shellcode on live hardware.
The exploit targets CVE-2021-31...
Source: The Hacker News (Security) | 6 hours ago
HuggingFace & Models
2. BenchMIRT: What are LLM benchmarks actually measuring?
Source: HuggingFace Blog | 16 hours ago
Research & Papers
3. HiveTraceGuard-Pro: A Compact Generative Guardrail for Prompt Injection, Jailbreaks, and Adversarial Obfuscation
arXiv:2609.01046v1 Announce Type: cross Abstract: Production LLMs must handle inputs that attempt to override system instructions, bypass safety policies or elicit harmful responses. A common mitigation is a separate guardrail model. Existing reports, however, provide little evidence on Russian ...
Source: arXiv - AI | 10 hours ago
4. Will the User Ever Know? Covert Indirect Prompt Injection Attacks on Tool-Using LLM Agents
arXiv:2608.30362v2 Announce Type: replace Abstract: As LLM agents take real-world actions through tools, indirect prompt injection (IPI) has emerged as a serious threat. The standard metric, Attack Success Rate (ASR), counts whether an injection succeeds but ignores what the user notices in the ...
Source: arXiv - AI | 10 hours ago
5. TACS: Trajectory-Aware Candidate Selection for LLM Jailbreak Suffix Optimization
arXiv:2608.29564v2 Announce Type: replace-cross Abstract: Gradient-based jailbreak suffix optimization methods typically update the suffix by retaining the candidate with the lowest current loss. We show that this seemingly natural design is fundamentally myopic: candidates that look better unde...
Source: arXiv - Machine Learning | 10 hours ago
6. Validity-Aware Jailbreak Evaluation for Large Language Models
arXiv:2609.00498v1 Announce Type: new Abstract: Jailbreak robustness has become central to large language model (LLM) safety evaluation, yet prevailing methodologies rely primarily on refusal behavior, semantic resemblance, and intent-matching heuristics that emphasize linguistic plausibility ra...
Source: arXiv - AI | 10 hours ago
7. On the Existence of Consistent Adversarial Attacks in High-Dimensional Linear Classification
arXiv:2506.12454v2 Announce Type: replace-cross Abstract: What fundamentally distinguishes an adversarial attack from a misclassification due to limited model expressivity or finite data? In this work, we investigate this question in the setting of high-dimensional binary classification, where s...
Source: arXiv - Machine Learning | 10 hours ago
Tech & Development
8. Why I prefer using LLM API aggregators over subscription services
Article URL: https://crazysteve.bearblog.dev/why-i-prefer-using-llm-api-aggregators-over-subscription-services/ Comments URL: https://news.ycombinator.com/item?id=49535211 Points: 2
Comments: 0
Source: Hacker News - AI | 1 hours ago
About This Digest
This digest is automatically curated from leading AI and tech news sources, filtered for relevance to AI security and the ML ecosystem. Stories are scored and ranked based on their relevance to model security, supply chain safety, and the broader AI landscape.
Want to see how your favorite models score on security? Check our model dashboard for trust scores on the top 500 HuggingFace models.