File Attack Deep Dives
File Attack Deep Dives: HTML as an Attack Surface

File Attack Deep Dives: HTML as an Attack Surface

Every way an HTML page attacks: smuggling, external fetches, QR handoffs, phone callbacks, credential phishing, job scams, and AI prompt injection.

HTML is the one file type that is also a program. It can fetch, decode, draw a convincing page, hand off to a phone, or talk to an AI — all after every transit control has already passed it. Judge the page by what it does, not what it looks like.

Part 1 of this series, the field guide, mapped the file attack families. This part takes a single file type — HTML — and walks every way it attacks: smuggling, external fetches, QR handoffs, phone callbacks, credential phishing, job scams, and AI-era prompt injection.

Why HTML is different

Most files carry a payload. HTML is the payload and the delivery at once: a document a gateway happily forwards, and a program a browser eagerly runs. In the same page it can decode an embedded file, reach out to a remote host, draw a pixel-perfect login screen, hand off to a phone, or carry text only an AI will read. Every vector below is one of those capabilities.

How does HTML smuggling land?

The delivery is an HTML page or HTML attachment. Inside it sits the payload as encoded text, usually Base64, plus a short script. When the victim opens the page, the browser follows the recipe and writes the real file to disk. The canonical flow, documented as MITRE T1027.006:

  1. The page carries the payload as an encoded string, often Base64 inside a JavaScript block.
  2. The script decodes it, typically with atob(), into raw bytes.
  3. It wraps the bytes in a Blob, an in-memory file-like object.
  4. It calls URL.createObjectURL(blob) to mint a temporary blob: URL.
  5. It creates an anchor element, sets href to the Blob URL and the download attribute to the chosen filename, and clicks it programmatically.
  6. The browser treats it as a download and saves the assembled file.

The whole sequence fits in a few lines:

const bytes = Uint8Array.from(atob(encodedPayload), (c) => c.charCodeAt(0));
const blob = new Blob([bytes], { type: "application/octet-stream" });
const a = document.createElement("a");
a.href = URL.createObjectURL(blob);
a.download = "invoice.iso";
a.click();

On older Internet Explorer and legacy Edge builds, the same outcome goes through window.navigator.msSaveBlob(blob, filename) instead of the anchor flow, so smuggling kits probe for it as a fallback path.

Strip the script away and what is left is a base64 string. Paste a sample into the Base64 workbench and you can decode it by hand to see the file it rebuilds. That is the whole trick: the malware is a text file until the browser turns it into a binary.

Step through the whole sequence and watch the gap open between what a control sees and what the machine actually does:

HTML smuggling, end to end

invoice.html

Delivered

What the control sees

text/html and a long Base64 blob — no executable, no archive

What actually happens

Nothing has run. The payload is inert text sitting in the page.

<script>
  const p = "TVqQAAMAAAAEAAAA//8AALgA…";  // truncated, defanged
</script>
Stage 1 of 4

The other ways an HTML page attacks

It fetches the payload for you

The page is a decoy — a "shared document", a delivery notice, an invoice. Its button, an image, or an automatic fetch pulls the real payload from a remote host the moment the page opens. The gateway saw a link, not the file; the file never existed in the attachment at all.

hxxp://acme-share[.]example/doc/8f21a9
Defanged example page titled 'A document has been shared with you' with an Open document button and an hxxp URL.
A decoy share page. The real payload is one outbound request away, triggered on open.
  • Tell: an HTML attachment whose only purpose is a single outbound request; a link that resolves to a raw executable or archive; content that appears only after the page loads.
  • Detection: detonate the page and follow the outbound chain; alert on HTML that downloads and then writes a file; apply URL category and reputation at the proxy, not just at the mail gateway.

It hands off to your phone

The payload is not the page — it is a QR code drawn inside it. Text scanners see a picture and move on. The human scans it with a phone, which leaves the corporate network, its proxy, and its logging entirely, and lands on a page built for that device.

hxxp://acme-verify[.]example/scan
Defanged example invoice page with a QR code and a 'Scan to verify and download' prompt.
An "invoice" with a QR handoff: the attack continues off-network, on the phone.
  • Tell: an image-heavy or image-only HTML page, a "scan to verify" prompt, or a QR whose decoded target is a bare IP, a shortener, or a look-alike domain.
  • Detection: decode QR images found in mail and HTML with the QR workbench or a resolver, then scan the target; treat QR-in-email as a callback vector, because by design it bypasses the desktop.

It asks you to call

No payload, no link — just manufactured urgency and a phone number. The attack completes on a call, where a person is talked into installing "support software", reading back an MFA code, or moving money. It is the one vector where the endpoint control is a human being.

hxxp://acme-security[.]example/alert
Defanged example security-alert page telling the reader to call a phone number and quoting a reference ID.
Callback phishing ("vishing"): the page manufactures a problem and a number to call.
  • Tell: a call-to-action that is a phone number, urgency, a "reference ID", and a request to keep the issue confidential.
  • Detection: this is social, not signature-based — flag callback and urgency language in inbound mail and HTML, and make the real control a rule people know: verify through a number you already have, never one the message supplies.

It impersonates a brand to steal credentials

The page is a pixel-accurate sign-in screen for a brand people trust — an identity provider, a bank, a payroll portal — sitting on a look-alike domain. Submitting the form harvests the password and, increasingly, the MFA code that follows it. Brand impersonation is not a separate trick here; it is the trust layer that makes every other HTML vector work.

hxxp://acme-sso[.]example/login
Defanged example single-sign-on login page titled Sign in to continue with email and password fields.
A cloned SSO screen on a look-alike domain. The form is the payload.
  • Tell: a login form inside an HTML attachment or on a non-canonical domain; credentials POSTing to an action unrelated to the brand; a domain that is brand-adjacent but wrong.
  • Detection: detonate the page and capture where the form actually submits; weigh domain age and DMARC alignment; the human tell is a password manager that refuses to autofill a domain it has never seen.

It offers you a job

The trust here is the offer, not a logo. A fake recruiter sends an "offer letter" page or document; the "sign and return" link harvests credentials, or the "onboarding app" is the malware. Job and HR themes work because the target is motivated and the urgency feels legitimate.

hxxp://acme-careers[.]example/offer/a-2291
Defanged example careers page showing an offer of employment with a salary, an attachment, and a 24-hour deadline.
A recruitment scam: a plausible offer plus a 24-hour deadline to click.
  • Tell: unsolicited offers, urgency ("sign within 24 hours"), a link to "verify" or "sign", and early requests for personal or bank details.
  • Detection: process over signature — verify the recruiter and the company through a channel you initiate, and never treat an offer attachment as safe because the story is convincing.

It talks to your AI assistant

The newest target is not the user or the endpoint — it is the model. Text that is invisible to a human (white on white, moved off-canvas, buried in metadata) is fully visible to the AI asked to read or summarise the document. The injected instruction runs with the assistant's access: the document stops being data and becomes a command.

hxxp://acme-billing[.]example/invoice/inv-2049
Defanged example invoice with a highlighted box revealing hidden white-on-white text that instructs an AI assistant to exfiltrate data.
Hidden text, revealed. A human sees an invoice; the assistant sees an instruction.
  • Tell: near-invisible text, instructions addressed to an AI ("ignore previous instructions"), exfiltration addresses, and documents routed into AI tooling with no extraction step.
  • Detection: extract and score documents before they reach a model, strip and flag hidden or anomalous text, and treat everything retrieved from a document as untrusted data — never as instructions.

Why do gateways miss all of it?

A gateway inspects what crosses the wire. In every vector above, what crosses the wire is benign-looking HTML and JavaScript. There is no executable, no archive, no container to detonate at the choke point. Microsoft's writeup of the technique puts it plainly: gateways only see benign HTML and JavaScript traffic while the malicious file is created on the endpoint after the page loads (HTML smuggling surges).

Three properties combine to make that gap reliable:

  • The action happens after inspection. Decoding, fetching, and rendering run at open time, on the victim's machine, past every transit scanner.
  • The page looks ordinary. Encoded strings, Blob calls, forms, and QR images are all things legitimate pages use, so naive pattern blocks drown in false positives.
  • The payload shape is free. Attackers split the encoded text across variables, nest encodings, swap the carrier for a QR or a link, or address the text to an AI, so no single static signature survives rotation.

This is not theoretical. In May 2024, Huntress documented a mass campaign that paired an HTML smuggling payload with an injected iframe proxying the Outlook login portal, stealing sessions when victims logged in (Smuggler's Gambit).

What does HTML-based detection look like?

Detection lives where the page acts: the endpoint, the browser, the phone, and the model. Work through these in order of signal quality.

  • Script pattern in the page. atob() or equivalent decoding feeding new Blob(), followed by URL.createObjectURL() and a programmatic .click() on an anchor with a download attribute. Any one call is common; the chained sequence in a single page is the tell.
  • Outbound fetch on open. An HTML attachment that reaches out to a remote host as soon as it renders — especially to a raw file, a bare IP, or a brand look-alike. Follow the chain, do not stop at the attachment.
  • QR images in mail or HTML. Decode them and scan the target. A QR that points off-network is a handoff, not a picture.
  • A form that submits somewhere unexpected. Capture the real action target of any login form and compare it to the brand it claims to be.
  • Browser-as-parent file writes. A browser process writing an executable, archive, or disk image, followed within seconds by execution, is the behavioral pair that survives obfuscation.
  • Hidden text in documents sent to AI tooling. Near-invisible or instruction-shaped text, and exfiltration addresses, extracted before the document reaches a model.
  • Missing Mark of the Web. A payload that runs without the Zone.Identifier marker it should carry suggests a container stripped provenance on the way in. Treat absent MOTW on an internet-born file as an anomaly, not a detail.

Pick a vector below, see how it lands, then reveal the tell and the detection.

HTML smuggling

How it lands. An HTML page or attachment carries Base64-encoded bytes plus script. On open, the browser decodes them, builds a Blob, mints a blob: URL, and triggers a download with the anchor download attribute.

How do you defend against it?

The fix follows the tell, and the same shape works for every vector: detonate the page and watch what it does, not just what it says. Run HTML attachments in a sandbox that executes script and fetches; follow the outbound chain; capture form actions; decode QR images; and extract documents before they reach a model. Block external HTML attachments outright where the business allows it — the blunt version of the same idea.

Does blocking HTML attachments stop all of this?

The email path, mostly. No HTML attachment means no in-browser assembly from mail. But the same techniques arrive via a link to a hosted page, a QR code, or a document handed to an AI, so endpoint and model-side detection still matter.

Why not just alert on every Blob and download call?

Because legitimate web apps build files the same way: exports, previews, and generated reports all use Blob plus the download attribute. Alert on the full chain — decoding into Blob construction into a programmatic download, ideally paired with a file write the endpoint did not fetch.

Is the AI prompt-injection vector a file problem or a model problem?

Both. The file is the delivery, the model is the target. You need extraction and scoring on the way in, and you need the assistant to treat document content as untrusted data rather than instructions. Neither side alone is enough.

Test yourself

A quick self-check on HTML-based attacks. Pick an option to see the answer.

0/4 answered · 0 correct
  1. 1. Where is the malicious file assembled in HTML smuggling?

  2. 2. Which API sequence is the core smuggling tell?

  3. 3. Why is a QR code in an HTML or email attachment hard for a gateway to catch?

  4. 4. What makes AI prompt injection in a document different from the other vectors?

What to remember

  • HTML is a program. It can fetch, decode, render a convincing page, hand off to a phone, or address a model — all after the gateway has passed it.
  • The tell is behavioral: the chained script pattern, an outbound fetch on open, a form posting somewhere unexpected, hidden text aimed at an AI.
  • Detection lives where the page acts — the endpoint, the browser, the phone, and the model — not just where the file travels.

Explore next:

Next in this series: one file type at a time, with every way it attacks — document macros, then PDF URL actions, then LNK and ISO containers.

Useful references:

Written by

Mohit Khare

Senior software lead at Abnormal.AI, prev. CRED and Gojek. Backend engineer working on AI, security, and engineering leadership. Based in Bangalore.

Connect on LinkedInMore writing