How the Canadian Centre for Cyber Security uses frontier AI to accelerate malware reverse engineering

An introduction to the Communications Security Establishment Canada's work on artificial intelligence in cyber defence

This technical article is the second in a series on the work of the Frontier AI Lab. For an introduction to the Frontier AI Lab, see our first article: "How the Canadian Centre for Cyber Security used frontier AI to accelerate detection engineering." The series shares findings from the Lab's research, including potential applications of frontier AI in cyber defence, the limitations of these technologies, and considerations for their effective, secure, and responsible use.

This second article focuses on malware reverse engineering, the process of deconstructing malicious software without its original source code to help cyber defenders understand its functionality and purpose. As cyber threats continue to evolve and organizations process increasing amounts of data, AI may help support certain aspects of this work.

Using practical test cases, the article examines how the pilot used AI to help turn unknown malware samples into validated intelligence packages. The evaluation explored how AI could assist analysts with tasks such as triage, detection engineering and reporting. Throughout the pilot, human oversight remained essential: analysts guided investigations, validated findings and reviewed attribution assessments.

On this page

Overview: How the AI-powered malware reverse engineering pipeline works

During the pilot, the Cyber Centre used frontier AI models to build an automated malware reverse engineering pipeline. The pipeline was designed to help turn unknown malware samples into validated intelligence packages with minimal human intervention.

For the samples evaluated, the packages included analysis reports, MITRE ATT&CK mapping, YARA detection rules, network signatures and configuration extractors prepared for deployment in Assemblyline 4, the Cyber Centre's malware detection and analysis tool.

This pilot project equipped Mythos with the same tools a reverse engineer would use, such as:

  • a disassembler and decompiler
  • the Cyber Centre's corpus containing tens of thousands of YARA detection rules
  • .NET and Python decompilers
  • CPU emulation frameworks

In practical terms, Mythos coordinated several specialized tools and paused at key stages to check the evidence before moving forward. This helped ensure that the system's findings were supported by observable results and that any proposed detection worked as intended.

Mythos then drove these tools through a checkpointed orchestration layer that enforced two safeguards:

  1. every factual claim in a report must trace to a verifiable tool result.
  2. every detection artifact must be validated against the real sample and assessed across a broader corpus before it is written to disk, ensuring it reliably identifies the intended artifact while minimizing unwanted detections.

Figure 1: Full malware reverse engineering pipeline

Long description immediately follows
Full malware reverse engineering pipeline - Long description

The figure shows the automated reverse engineering workflow from sample ingestion through human review. It identifies the stages of the pipeline: classification and YARA corpus triage, routing to a native, .NET, or script analysis path, parallel deep analysis across six technical domains, payload extraction and recursive re-analysis of recovered stages, artifact generation, and validation. The figure demonstrates how deterministic validation gates surround the model at each stage before artifacts reach analyst review.

The pipeline begins by ingesting malware samples that are submitted directly to the Cyber Centre. The following test cases illustrate how this workflow was applied in practice.

Malware sample 1: Analysis of an obfuscated loader

The first malware sample demonstrates how the pipeline handled deep reverse-engineering complexity.

Malware

The first sample was designed to make its hidden content difficult to recover. It combined an unfamiliar encryption process with several layers of obfuscation, requiring the pipeline to reconstruct how the malware generated its keys before it could reveal the concealed payload.

This malware sample was a loader written in Go (Golang) and processed with an obfuscating compiler. It protected its payload using a scheme that its internal strings described as "PQC-PE-AES256GCM". The post-quantum key encapsulation algorithm (ML-KEM-768) was used to derive the keys for 48 AES-256-GCM encrypted chunks, while the underlying key material was further split, shuffled, and masked.

Note: In this context, post-quantum cryptography likely did not provide a meaningful operational security advantage to the malware. Its likely purpose was to increase analytical complexity by forcing defenders to understand and reproduce an unfamiliar key encapsulation and decryption workflow before the payload could be recovered.

YARA triage, classify and route

All samples are scanned against the Cyber Centre's full YARA corpus, which contains tens of thousands of compiled rules. This triage is applied not only to each incoming sample, but also to every recovered payload since downstream stages may be just as likely to match a known malware family as the file that delivered it.

This initial triage step is critical because a corpus match can fundamentally change the economics of the engagement and provides three immediate benefits:

  1. A working identity: a family-labelled hit connects the sample to everything the Cyber Centre already knows about that malware family. If an existing extractor can recover the command-and-control configuration, the engagement can end within minutes, allowing the sample to be routed through established tooling rather than requiring an analyst.
  2. Evidence-based starting points: every rule includes the offsets of the strings or byte patterns that fired. Since the rule documents the family features it was built to detect, those offsets provide concrete starting coordinates for verification and analysis.
  3. Useful negative evidence: even imperfect matches are informative. If a sample triggers a family rule but other evidence does not align, the discrepancy becomes a specific hypothesis to investigate.

Together, these outcomes help turn open-ended investigations into focused, evidence-driven analysis.

YARA triage helps transform early findings into focused hypotheses for testing, rather than requiring analysts to explore a sample from scratch. In this case, neither the original sample nor its recovered payloads matched any known malware families in the Cyber Centre's YARA corpus, so the workflow proceeded to full analysis.

AI deep analysis and payload extraction

To recover the hidden payload, the model first had to understand the malware's logic and then safely reproduce only the parts of its behaviour needed to generate the correct decryption keys. The following steps describe how it combined code analysis and controlled emulation to do this.

The model used the decompiler to work through the loader's obfuscated code. It identified the cryptographic scheme by analyzing the code structure rather than matching known constants, then reconstructed the key-derivation chain.

Static analysis eventually reached key material that the malware generated only while running. To address this, the agent wrote a custom central processing unit (CPU) emulation harness for the Go runtime. It then executed the required derivation logic and captured the resulting keys and chunk schedule.

Because the malware had been deliberately altered to resist analysis, it could not simply be run from beginning to end in the emulator. The system instead worked through failures one at a time, preserving the functions needed for decryption while safely simulating less important parts of the program.

Emulating an obfuscated Go binary is often an iterative, analyst-driven process. Each execution may fail in different runtime functions, requiring repeated cycles of crash analysis, functionality emulation, and re-execution. In this case, the agent automated this process by creating the following self-improving loop around its harness:

  • when a crash occurred, the agent used the disassembler to locate the function containing the faulting instruction, replaced it with a stub, then re-ran the sample
  • a protected function list prevented key derivation logic from being modified. If a crash occurred within a protected function, the agent treated it as evidence that an upstream caller had supplied invalid input and stubbed the calling function, as opposed to the target.
  • runtime functions with real semantics were implemented accurately rather than simplified stubs, as they can produce plausible but incorrect keys
  • every recorded value was independently re-derived in a separate Python implementation before being used by subsequent analysis stages
  • if the same function caused repeated crashes, the loop stopped and was escalated to a human analyst

Figure 2: The simplified self-improving loop

Long description immediately follows
The simplified self-improving loop - Long description

The figure depicts an iterative emulation workflow where a crash with a faulting address triggers a Disassembler lookup from the Unicorn harness (round N), producing either a new stub or a predecessor stub when the function lies on a protected key path. A decision stage then determines whether a capture fired—if so, keys and a chunk schedule are verified in pure Python; if not, the system checks whether the same function has crashed twice, in which case execution stops and is escalated to an analyst, otherwise the process advances to the next round.

This step mattered because the sample's key derivation depended on the outcome of its anti-debugging logic. Simply bypassing those checks allowed emulation to continue, but produced a plausible yet incorrect key.

During the pilot, the model and malware-agent harness identified the relationship between the anti-debugging results and the derived key material. The harness was then adjusted to preserve the required behaviour rather than bypass it. This approach was subsequently added to the evaluation pipeline as a reusable capability for samples whose anti-analysis mechanisms affect decryption or payload recovery.

The keys and chunk schedule recovered through emulation were then used to build a payload extractor for the loader. By reproducing the sample's decryption workflow outside the malware, the extractor statically recovered the embedded payload.

Proposed detections, validation, and output

During the evaluation, the extractor's output was compared with a payload previously recovered from live memory through manual analyst debugging. The 195 KB payload recovered statically was byte-identical to the manually recovered version.

Testing against sibling samples indicated that the technique worked across five distinct builds of the family. Each recovered stage was then fed back through the pipeline for re-scanning and analysis. This process identified the known malware family ACR Stealer and linked the novel loader to previously observed threat activity.

As a result, a series of deliverables were produced and reviewed by an analyst prior to being sent to a client or put into production:

  • a full technical report with MITRE ATT&CK mapping
  • YARA rules for the loader and its delivery artifacts
  • a configuration/payload extractor packaged for the Cyber Centre's automated malware analysis platform using the open source MACO framework

Figure 3: Gatestomp payload extractor results

Long description immediately follows
Gatestomp payload extractor results - Long description

The image shows a screenshot of a "ConfigExtractor" interface displaying details for the GateStomp loader and a description referencing a PQC-PE-AES256GCM configuration parser. Below, a JSON-like section labeled "Other data" lists a payload under "binaries" with encoded data and "decoded_strings" that include Windows system file paths such as ntdll.dll and wmsxml3.dll. Overall, the interface illustrates how the analyzer surfaces both high‑level attributes and low‑level artifacts to help reverse engineers validate the family, confirm encryption and packing characteristics, and pivot on decoded indicators during threat research.

Malware sample 2: An unknown, analysis-resistant installer

The second malware sample demonstrates the pipeline's ability to quickly recover the delivery chain, avoiding distractions from the sample's size, signatures, and decoy content.

Malware

This malware sample involved three oversized, validly signed installer packages that were designed to look like ordinary commercial software and waste automated analysis effort. Each package:

  • was between 76 MB and 175 MB
  • carried a genuine Microsoft Authenticode signature
  • contained legitimate-looking structure, embedded runtimes, and large volumes of decoy content

The malicious functionality was small but deliberately buried.

YARA triage, classify and route

As previously stated, all samples are scanned against the Cyber Centre's full YARA corpus, which contains tens of thousands of compiled rules. This triage is applied not only to each incoming sample, but also to every recovered payload since downstream stages may be just as likely to match a known malware family as the file that delivered it.

This initial triage step is critical because a corpus match can fundamentally change the economics of the engagement and provides three immediate benefits:

  1. A working identity: A family-labelled hit connects the sample to everything the Cyber Centre already knows about that malware family. If an existing extractor can recover the command-and-control configuration, the engagement can end within minutes, allowing the sample to be routed through established tooling rather than requiring an analyst.
  2. Evidence-based starting points: Every rule includes the offsets of the strings or byte patterns that fired. Since the rule documents the family features it was built to detect, those offsets provide concrete starting coordinates for verification and analysis.
  3. Useful negative evidence: Even imperfect matches are informative. If a sample triggers a family rule but other evidence does not align, the discrepancy becomes a specific hypothesis to investigate.

Together, these outcomes help turn open-ended investigations into focused, evidence-driven analysis.

Initial triage identified the files as OLE compound documents and MSI installers, not conventional executables. With no malware family match across the full YARA corpus, the sample was deemed unknown and subsequently routed for full analysis.

AI deep analysis and payload extraction

For the second sample, the main task was to separate the small amount of malicious code from a large volume of legitimate-looking files and filler. Rather than examining every file in depth, the pipeline followed the installer's structure to identify which components were actually executed.

The analysis recovered the delivery chain statically. Each MSI package, built with a commodity wrapping tool, carried a cabinet archive inside an OLE stream. Expanding the archive produced 105 to 144 files that were between 137 MB and 231 MB uncompressed, almost all of which were decoys, and included:

  • a one-line launcher script
  • a complete and legitimate embedded Python runtime
  • an entry script disguised as an innocuous text file

The entry script was the family's defining feature. Buried within pages of textual padding was a single line that decoded, first from Base64 and then from UTF-32, into an eight-line downloader.

The downloader fetched content from a URL over a Transport Layer Security (TLS) channel without verifying the certificate, waited for a fixed interval and executed the response. This short sequence became the focus of the evaluation because it revealed the sample's core malicious behaviour.

The actor invested in:

  • a purchased certificate
  • a 100-file decoy corpus
  • hundreds of megabytes of padding
  • an embedded runtime (designed so that neither automated tools nor human analysts could inspect closely)

The challenge was not the complexity of the delivery chain, but the volume of content surrounding it. The model followed the structure of the package directly from the:

  • installer to the archive
  • archive to the launcher
  • launcher to the file it executed

Once there, the hidden line stood out because it was the only content that decoded into a meaningful payload. In this evaluation, the workflow completed the structural walk and decoding process in under one second per sample.

Figure 4: One line among the decoys

Long description immediately follows
One line among the decoys - Long description

The figure shows, on the left, an illustrative reconstruction of the family's entry script — pages of dictionary-word padding surrounding one highlighted encoded line — and, on the right, the behaviour of the eight-line downloader that line decodes to, with the code itself withheld: standard-library imports, disabled TLS certificate verification, a fetch from attacker infrastructure, a fixed delay, and execution of the response. A caption notes the defender's asymmetry: locating the line manually requires expanding a multi-hundred-megabyte archive, while the extractor recovers and decodes it structurally in milliseconds.

Each recovered stage was re-scanned against the Cyber Centre's corpus. However, the command-and-control (C2) infrastructure was no longer live, preventing observation of the server-fetched stage. Rather than speculate, the report recorded it as unknown.

Validate, analyst review, and production

Once the delivery chain had been recovered, the next challenge was to identify stable characteristics that defenders could reliably detect. The analysis therefore focused on behaviours that were essential to how the malware operated, rather than details the developer could easily change between versions.

Comparison of the three samples showed that the builder randomized nearly every feature an analyst might use for attribution. The only consistent artifact across all builds was the phrase "decoy padding", earning the name Lexiloader, derived from "lexicon". However, the detection logic did not rely on the name or the decoy padding alone. It anchored on traits that cannot be rotated per build without re-engineering the delivery chain itself, such as:

  • the download-and-execute construction
  • the deliberate disabling of TLS verification
  • the text-decoding sequence that turns one padded line into running code

Approximately three hours after intake, the pilot pipeline produced a series of deliverables. Producing the same set manually would normally require substantial analyst effort:

  • a full configuration extractor staged in the Assemblyline 4 automated malware analysis framework, requiring only a standard library parse of the OLE structure, one archive expansion, and one decode
  • automated checks using known-correct results to confirm that the extractor consistently recovered the same configuration and payload from each sample
  • validation against all three samples using golden reference outputs, emitting both the extracted C2 configuration and the decoded stager as child artifacts for automatic re-scanning in Assemblyline
  • three candidate YARA rules covering the carrier, the stager, and an observed carrier-agnostic variant, each passing the Cyber Centre's YARA rule validator without errors or warnings
  • a full technical report with MITRE ATT&CK mapping
  • YARA rules for the installer carrier, decoded stager, and carrier-agnostic variant
  • a configuration/payload extractor packaged for the Cyber Centre's automated malware analysis platform using the open source MACO framework

Figure 5: Lexiloader extractor results

Long description immediately follows
Lexiloader extractor results - Long description

This figure shows the output of a Lexiloader configuration extractor within the Assemblyline malware analysis platform. The extractor identified the malware family, recovered command-and-control network indicators, and decoded an embedded stage-1 Python downloader, displaying the associated URL, hostname, communication details, and extracted payload metadata. The recovered configuration is presented in a structured format, including the decoded payload, extraction method, cryptographic hash, and details of the loader's staging process. The results demonstrate the successful automated recovery and decoding of malware configuration data for further analysis and detection.

The rules remain in testing pending promotion. In this evaluation, roughly five agent-hours produced work estimated to be equivalent to approximately one week of conventional analyst effort. The initial structural assessment—that the samples did not match a family in the corpus and identifying their delivery chain—was completed through six minutes of deterministic checks at intake.

Results and operational benefits

Manual reverse engineering at this depth can require several days of analyst effort, or weeks for complex samples. During the evaluation period, the pipeline completed five complex engagements across four distinct malware families in approximately 46 hours.

The evaluation showed the clearest efficiencies when the malware-agent harness reduced tasks that might otherwise require hours of specialist work to minutes of tool-assisted analysis. These results were observed across several different samples. Other pilot highlights included:

  • Malware sample 3: full recovery of a five-stage PowerShell attack chain without executing attacker code. The model reduced 217 KB of heavily obfuscated code to 15 KB of readable source, enabling static recovery of later stages that would have been missed by traditional sandbox analysis.
  • Malware sample 4: linkage of two steganographic loader samples to the same builder run through byte-level analysis. Both samples were fully unwound across a four-stage decode chain, recovering all 31 internal loader modules and the payload's complete command-and-control configurations.
  • Malware sample 5: identifying a previously unnamed malware family within six minutes, after four independent structural checks contradicted the label accompanying the samples. The checks were deliberately cheap and deterministic:
    • file-format identification (signed MSI files)
    • a nine-byte code-signature gate in the existing family's configuration extractor
    • bi-directional YARA testing the labelled family's rules against carriers and unpacked payloads, and vice versa
    • a sandbox-behaviour anomaly, flagged at intake by the analyst

Overall, the pilot suggests that frontier AI can help accelerate parts of malware analysis and support the production of timely, high-quality intelligence. The evaluation also showed that the most important components were the safeguards around the model, designed and guided by experienced analysts. By automating selected, mechanical parts of the reverse engineering workflow, the pilot allowed analysts to spend less time on repetitive tasks and more time:

  • creating, reviewing and approving intelligence products
  • pursuing campaign-level questions often deferred by operational demands
  • converting analytical lessons into lasting pipeline improvements

Next steps for the Cyber Centre's use of frontier AI

The capability continues to mature and lessons from each investigation are being captured in a reusable agent skill library. This allows improvements identified during one investigation to strengthen the pipeline's performance in subsequent investigations.

Additionally, validated configuration and payload extractors are packaged for use in Assemblyline 4, allowing for automated payload analysis through proven malware-analysis workflows.

The Cyber Centre will continue sharing findings with the broader cyber defender community through established partnerships and open-source contributions.

Date modified: