Skip to main content
← Back to Blog
Security & DevOps8 min read

The ARTEX Bank Hacks Turned Agent Memory Into Forensic Evidence

CrowdStrike says ARTEX and multiple LLMs aided attacks on South Korean banks. Exposed agent memory then revealed the operation’s methods and limits in detail.

A financially motivated attacker used an AI penetration-testing agent and several large language models during a campaign against South Korean financial organizations, according to an October 7 investigation by CrowdStrike. The security company says the activity resulted in stolen data and unfolded from late September into early October. At least nine banks have disclosed attacks or been identified in Korean media reports, although authorities have not established that every incident came from the same operator.

The campaign matters because the evidence goes beyond a screenshot of a chatbot or a vague claim that “AI was involved.” CrowdStrike found exposed directories containing ARTEX configuration, coding-agent session histories, prompts, and agent memory files. Those artifacts reportedly documented a two-server operation, the models involved, parts of the target list, and the operator's workflow.

Persistent context appears on both sides of the incident. It helped the offensive system carry instructions and discoveries across sessions. When the attacker left that state exposed, the same memory became a forensic record. The lesson is not that memory is inherently dangerous. It is that agent memory is operational data: valuable, sensitive, attributable, and subject to the same retention and access-control failures as logs, credentials, and case notes.

What investigators found

CrowdStrike's report says one attacker-controlled server hosted ARTEX, an agentic penetration-testing framework released as open-source software earlier this year. A second server served as the primary operational environment. Open directories on that infrastructure contained a standing Chinese-language penetration-testing prompt, Claude Code session histories, ARTEX files, and Claude memory files.

ARTEX is not itself a foundation model. It orchestrates models and security tools to automate authorized penetration-testing work. CrowdStrike says the observed deployment used DeepSeek v4.1-flash as its principal model backend and supplemented it with GLM-5.3 and Grok 4.6 in additional coding-agent sessions. The company mapped the activity to standard MITRE ATT&CK techniques for acquiring infrastructure, obtaining AI capabilities, and using proxies.

The distinction matters. Describing this as “Claude hacked nine banks” would be inaccurate. The evidence points to a human operator assembling infrastructure, an orchestration framework, multiple models, proxy services, and familiar offensive techniques. AI increased the pace and reach of the workflow; it did not erase the operator, the vulnerable systems, or the conventional infrastructure that made the intrusions possible.

CrowdStrike attributes the campaign to no named threat group. It assesses with moderate confidence that the actor was Chinese-speaking and financially motivated, based partly on Chinese-language prompts and the origin of the tooling. Personal details appeared in one model session, but CrowdStrike explicitly says available evidence cannot definitively connect those details to the attacker. They are not repeated here.

The confirmed impact is narrower than the largest headlines

Reuters reported that South Korean police opened an investigation after attacks affected multiple banks. Shinhan Bank said information belonging to roughly 25,000 customers was compromised; KB Kookmin Bank reported exposure involving 119 customers. CrowdStrike said its observed targets overlap with organizations named in public reporting, but its report also says the total number of affected organizations remained unconfirmed.

That leaves several unresolved questions. Public evidence does not show that every reported bank incident was executed by ARTEX, that the agent autonomously chose targets, or that AI produced every successful exploit. It also does not establish the identity represented in the leaked sessions. The reliable claim is more specific: investigators found operational artifacts connecting agentic tooling and multiple LLMs to a campaign against South Korean financial organizations that resulted in data exfiltration.

On October 9, Reuters reported that ARTEX's developer had removed the public repository and said future development would continue as closed source because of misuse. Removing source code may reduce casual access to later versions, but copies of an already published project can persist indefinitely. It also makes independent auditing and defensive research harder. The security question is therefore larger than whether one repository remains online.

Agent memory changes the evidence trail

Traditional intrusion investigations assemble a story from network traffic, authentication records, malware, command histories, and files left on compromised systems. Agentic workflows add a new class of artifact. Session transcripts can reveal instructions and attempted actions. Configuration identifies models, tools, and permissions. Memory files preserve facts the system carried forward: targets, credentials, preferences, discoveries, failures, and next steps.

That persistence can make an operator faster. An agent does not need to rediscover the same endpoint or be briefed again after every restart. It can retain a target map, reuse successful procedures, and coordinate subtasks across a campaign. A lone person can therefore operate with some of the continuity previously associated with a larger team.

The same continuity increases exposure. If memory is stored in a readable directory, copied into prompts, synchronized to the wrong host, or retained without a deletion policy, a single mistake can disclose the history of an operation. Model conversations may also contain identifiers and personal details that an operator would never place in a purpose-built command log. Fluent tools can make unsafe disclosure feel like ordinary conversation.

Defenders should not assume those artifacts will always be available. A careful attacker can encrypt them, isolate them, rotate infrastructure, or disable persistence. This campaign is valuable precisely because the exposed directories offered an unusually clear view. It demonstrates what agent telemetry can reveal, not a guaranteed method for attributing future AI-assisted attacks.

What responsible agent deployments should record

Organizations deploying agents for legitimate security work need a deliberate evidence policy. Records should identify the human or service principal that initiated a task, the model and tool versions used, the authorization scope, tool calls, data destinations, and material state changes. Logs should be tamper-evident, encrypted, tenant-isolated, and retained only as long as a stated operational or legal purpose requires.

Memory deserves separate treatment from audit logs. An audit log should provide a durable account of what happened. Working memory should contain only the minimum context needed for the current task and should not silently become an unlimited archive. Sensitive values should be referenced through scoped secret stores rather than copied into conversational history. Superseded targets, expired credentials, and completed cases should leave active memory without erasing the independently protected audit record.

Tool permissions also need to be enforced outside the model. A warning in a prompt does not prevent scanning, exploitation, or exfiltration. Network policy, sandboxing, identity controls, destination allowlists, rate limits, and human approval gates create boundaries that remain effective when a model misunderstands or disregards instructions. The operator in this case combined AI with ordinary infrastructure; defenders need controls at those ordinary layers.

The shift is operational, not magical

This incident does not prove that autonomous systems can compromise any target without expertise. It does show that commercially available models and an orchestration layer can package reconnaissance and exploitation into a reusable workflow. That lowers the labor required to run several operations, while preserving enough state for investigators to reconstruct the process when the operator makes a basic security mistake.

The durable security insight is almost mundane: an agent's memory is a database with consequences. Offensive operators can use it to scale. Defenders can use it for accountability. Anyone deploying agents should decide, before the first task runs, what gets remembered, who can read it, how contradictions and corrections are recorded, when working state expires, and which audit evidence must survive. Otherwise the system will make those choices accidentally.

Sources

Your AI remembers everything. Everywhere.

Unimatrix gives you a shared, durable memory layer across Claude Desktop, Cursor, ChatGPT, and Gemini. Setup in 2 minutes. Free and paid plans available.

Looking for developer resources?

Browse our catalog of 500+ tested AI prompt profiles covering DevOps, data modeling, agent behaviors, and API wrappers. Have a prompt to share? Submit your own for review by our librarian to be featured. Completely free, no registration required. Browse prompt libraries →

Keep reading