← Writing
Parsing

The edge cases I hit building a live-preview Markdown editor

I assumed Markdown would be the easy part: grab a CommonMark parser and move on to the interesting problems. The format had other plans.

I've been building a Markdown notes app on weekends. It's called Margin, the notes live as plain .md files on disk, and I assumed Markdown itself would be the easy part: grab a CommonMark parser and move on to the interesting problems.

The format had other plans. Four things ate whole weekends, and none of them were the ones I expected.

Pipes in table cells

A | inside a table cell has to be escaped, or it splits the cell. That much is obvious. What isn't obvious is what happens inside an inline code span.

GFM splits a table row into cells before inline parsing runs. So the code span doesn't exist yet at the point the pipe is interpreted, which means a literal pipe inside a code span in a table still has to be written \|.

The backslash gets eaten by the table parser, not by the code span.

Everyone I've watched hit this assumes the code span protects the pipe, which is an entirely reasonable thing to assume. It just isn't how the order of operations works.

That leaves the serializer with a job that sounds trivial and isn't: escape every pipe on the way out, unescape it on the way back in, and make sure a round trip through both lands exactly where it started.

src/lib/mdTable.tsts
// Split one table line into trimmed, unescaped cells. Outer pipes are
// optional; a `\|` inside a cell is a literal pipe, not a separator.
function splitCells(line: string): string[] {
  let s = line.trim();
  if (s.startsWith("|")) s = s.slice(1);
  if (/(?:^|[^\\])\|$/.test(s)) s = s.replace(/\|$/, "");
  return s
    .split(/(?<!\\)\|/)
    .map((c) => c.trim().replace(/\\\|/g, "|"));
}

Getting that function and its opposite number to be genuine inverses of each other took embarrassingly long.

Round-tripping

This is the one that changed how I think about the whole app.

A live-preview editor parses your text, renders it, and writes it back. The write-back has to reproduce exactly what you typed. If you write * bullets and the serializer prefers -, the app has silently rewritten your file. If you indent with three spaces and it normalises to two, same thing.

Because the files are plain Markdown that other tools might read, and because the entire promise of the app is that your notes are just files, I decided to treat any normalisation as data corruption. Not as a cosmetic issue, as corruption.

Holding that line is harder than it sounds, because normalising is the natural thing for a serializer to do. Every convenience you add pulls toward a canonical form.

The same principle shows up in how the table parser fails. If it gets something that isn't a well-formed table, it doesn't try to repair it. It returns nothing, and the caller falls back to rendering the source as plain text. A guessed repair means a slightly-malformed table silently becomes a different table the first time it round-trips, and the original is gone. Refusing to parse is the safer failure.

Paste

Everything arrives as HTML.

Paste a table out of a Google Doc and you get nested span soup with inline styles on every cell. Paste from a webpage and you get smart quotes, non-breaking spaces, and markup that means nothing once it's Markdown. Converting that into Markdown which then survives a round trip through the previous section's rules is its own project, and I'm still not finished with it.

The cursor

In live preview, **bold** renders as bold until your cursor enters it, at which point the markers reveal themselves so you can edit the raw text.

Exactly when they appear, and where the caret lands when they do, is a product decision wearing a parsing costume. Reveal too eagerly and the text jumps around while you type. Reveal too late and you can't get at your own syntax. Put the caret one character off after a reveal and the next keystroke lands inside the markers instead of outside them.

You feel every wrong choice within minutes of typing, which is a good property for a problem to have. It's also why I ended up on CodeMirror 6 rather than a contenteditable surface. The document is always plain text and bold is just ** characters, so there's no formatting state at the cursor for the browser to extend into. The rendered look is decorations painted on top, and the browser never owns the editing.

There's no HTML render step anywhere in the app, which makes "which implementation do you render with" a slightly odd question to answer. The answer is none of them.


If I'd known at the start that the parser was the easy half, I'd have budgeted differently. The hard half is everything that happens after you have a correct parse: putting the text back exactly as it was, deciding what not to touch, and working out where a blinking rectangle belongs.

Margin is the app this came out of
Local-first Markdown notes for the Mac. Free while it is alpha.
Download the alpha
More writing
Parsing
Three things I decided not to parse

Unmatched embeds, tables that refuse to be repaired, and display options with nowhere to live. Sometimes declining to parse is the safer failure.

16 Aug 2026·2 min·/blog/not-to-parse
On disk
The watcher that cried delete

Some editors save by replacing the whole file, so a file watcher briefly sees a deletion. Believe that event and your app tells someone their note is gone.

8 Aug 2026·3 min·/blog/atomic-saves
Building
16 lessons from building a productivity app on weekends

Notes from the workshop rather than wisdom from the mountain. What a few months of building a notes app has changed about how I look at every app I use.

2 Aug 2026·4 min·/blog/sixteen-lessons