mailfile
v0.1.1
Published
Read and write Outlook .msg and RFC 5322 .eml files. Zero dependencies, runs in browsers, Node, Deno, Bun and workers.
Maintainers
Readme
mailfile
Read and write Outlook .msg and RFC 5322 .eml files. Zero dependencies,
one small model in the middle, and it runs wherever TextDecoder does —
browsers, Node, Deno, Bun, Cloudflare Workers.
.msg ──► [cfb] ──► [mapi] ──┐ ┌──► [mapi] ──► [cfb] ──► .msg
├──► Message ───┤
.eml ──────────► [mime] ────┘ └──────────────[mime] ──► .eml
▲
[rtf] recovers the HTML body when a .msg has no PR_HTMLTwo parsers in, two serializers out. The Message between them is plain and
mutable, so this is a toolkit as much as a converter: extract attachments,
rewrite headers in bulk, redact recipients, index a mail archive.
Install
npm install mailfileOr drop it straight into a page — no build step, no bundler:
<script src="https://cdn.jsdelivr.net/npm/mailfile@0/dist/mailfile.min.js"></script>
<script>
const { data, format } = mailfile.convert(bytes, { filename: 'note.msg' })
</script>Convert
The common case is one call. The input format is detected from its magic
bytes, so a .msg someone renamed to .eml is still handled correctly.
import { convert } from 'mailfile'
const { data, format, message } = convert(bytes, { filename: 'report.msg' })
// data -> Uint8Array of .eml bytes
// format -> 'eml'
// message.subject, message.attachments, ...Pass to when you want a specific format rather than the other one:
convert(bytes, { to: 'eml' }) // normalise anything to .emlThe message model
import { Message } from 'mailfile'
const m = Message.fromMsg(bytes) // or Message.fromEml(bytes)
m.subject // 'Quarterly résumé — café'
m.from // { name: 'Alice Åberg', email: '[email protected]' }
m.to // [{ name, email }, ...] also .cc and .bcc
m.date // Date | null
m.text // plain-text body
m.html // HTML body, recovered from compressed RTF if needed
m.attachments // [{ filename, mime, data, cid, inline }]
m.headers // ordered header multimap, see below
m.subject = 'Re: ' + m.subject
m.attachments = m.attachments.filter(a => !a.inline)
const eml = m.toEml()
const msg = m.toMsg()Headers keep their order and their repeats:
m.headers.get('subject') // first value, case-insensitive
m.headers.getAll('received') // every hop, in order
m.headers.set('X-Reviewed', 'yes')
m.headers.raw // the block as text, foldedOriginal internet headers survive a .msg → .eml conversion when the
message carried them, so a .msg that started life as email keeps its full
Received: trail. Content headers are always regenerated to match the body
actually written.
Reading only what you need
lazy defers pulling attachment bytes out of the compound file until you
touch .data — worth it when you are indexing headers across a large archive.
const m = Message.fromMsg(bytes, { lazy: true })
m.subject // cheap
m.attachments[0].data // extracted nowProperties the model does not cover
.msg is a MAPI property bag, and this library models the mail-shaped subset.
Anything else is still reachable — appointments, contacts, custom properties:
const m = Message.fromMsg(bytes)
m.messageClass // 'IPM.Appointment'
m.props.str(0x0037) // PR_SUBJECT
m.props.int(0x0E07) // PR_MESSAGE_FLAGS
m.props.date(0x0060) // PR_START_DATE
m.props.ids() // every property id presentThe pieces, separately
Each layer is independently useful and independently importable, so you only pay for what you use.
import { read, write, newRoot, addStream } from 'mailfile/cfb'
import { parseEml, buildEml, Headers } from 'mailfile/mime'
import { PropertyBag, PropId } from 'mailfile/mapi'
import { decompress, rtfToHtml } from 'mailfile/rtf'mailfile/cfb is a complete OLE2 compound-file reader and writer. It has
nothing to do with email — point it at a legacy .doc, .xls, .ppt, an
.msi, or a Thumbs.db:
const cfb = read(bytes)
for (const entry of cfb.childrenOf(cfb.root).list) {
console.log(entry.name, entry.size)
}Errors
Every failure is a MailfileError carrying a code, so you can tell "wrong
kind of file" apart from "damaged file" without matching on message text.
import { MailfileError, ErrorCode } from 'mailfile'
try {
convert(bytes)
} catch (err) {
if (err.code === ErrorCode.NOT_COMPOUND_FILE) // not really a .msg
if (err.code === ErrorCode.CORRUPT_CFB) // damaged
}Parsing is deliberately forgiving — real mail is malformed constantly, and a mislabelled charset or a missing closing boundary should not cost you the message.
What carries across
Headers, sender and recipients, subject and date, plain-text and HTML bodies,
inline images by content ID, and every attachment byte for byte. HTML stored
only as compressed RTF (PR_RTF_COMPRESSED) is de-encapsulated back to HTML,
which is how a lot of Outlook-authored mail stores it.
Known gaps
- Named properties.
__nameid_version1.0is written empty and not parsed, so categories and custom named properties do not survive a write. - Embedded messages are read (they become nested
.emlattachments) but written back by value, not as true embedded messages. PR_RTF_COMPRESSEDis never generated. Outlook regenerates it on open; stricter MAPI consumers may not.- TNEF (
winmail.dat) is not handled. - PST/OST is out of scope and will stay that way.
Testing
npm test # unit tests
node scripts/corpus.js --eml ~/mail --any # your own mail, end to endThe corpus runner takes any directory tree of real messages, pushes each one
through eml → msg → eml, compares field by field, and has an unrelated MSG
reader check everything written. It is the check that finds real bugs; bring
your own archive, since none ships here.
Implementation notes
Written against the published Microsoft open specifications — MS-CFB
(compound file), MS-OXMSG (.msg), MS-OXRTFCP (RTF compression) and
MS-OXRTFEX (HTML encapsulation) — with no code derived from an existing
implementation.
License
MIT
