Skip to main content
45dd173c229b363e56d0
·7 min read

Moving from Shiki to TanStack Highlight

Why I replaced Shiki on this site, what changed in the code, and which TanStack Highlight features made the migration worthwhile.

tanstack
highlight
react
shiki
typescript
3ef4f885272bc6323209
Sohan R. Emon

Developer, Learner, Tech Enthusiast

I used Shiki for the code blocks on this site. It worked well. I liked the themes, and its transformers made it easy to mark diff lines and highlight selected lines.

After using it for a while, I noticed that the rendering path was more complicated than the feature needed. The site mostly shows documentation, and it already knows the language of every code block.

What I had with Shiki

The old code called codeToHtml from a TanStack Start server function. The browser waited for that function and then updated the code block after hydration:

ts
const highlightCode = createServerFn({ method: 'GET' }).handler(async ({ data }) => {
  return codeToHtml(data.code, {
    lang: data.lang,
    theme: 'one-dark-pro',
    transformers,
  })
})

The component kept the result in state and rendered it with dangerouslySetInnerHTML. I also used Shiki transformers to add classes to diff lines and selected lines.

It worked, but every code block had an asynchronous path. There was a server function, an idle hydration boundary, a loading state, and a fallback when highlighting failed.

What changed

TanStack Highlight creates a synchronous highlighter that I can share between the server and the browser:

ts
const highlighter = createHighlighter({
  fallbackLanguage: 'plaintext',
  languages: allLanguages,
})

The code is tokenised directly, and renderTokens returns a tree that my React component can render:

ts
const { tokens } = highlighter.tokenize(code, { lang })

const nodes = renderTokens(tokens, {
  lineNumbers: true,
  decorations,
})

This removed the server function, useEffect, loading state, idle hydration wrapper, and dangerouslySetInnerHTML. The code is rendered during SSR and stays the same after hydration.

The difference in this project

The comparison looks like this when I map it to the code I actually had:

AreaShiki beforeTanStack Highlight now
SetupAn async server function returned highlighted HTMLOne synchronous highlighter is reused at module scope
RenderingReact injected the HTML with dangerouslySetInnerHTMLReact renders the token tree returned by renderTokens
HydrationCode blocks waited for idle hydrationCode blocks are available during SSR
ThemesA fixed one-dark-pro theme produced the outputSemantic token classes let CSS handle light and dark themes
Highlighted linesShiki transformers changed the output treeLine and character-range decorations carry the extra state
Diff blocksCustom transformer logic added classes and signsDecorations add classes, while the component renders the signs
Language loadingShiki loaded its grammar and initialised before useTanStack Highlight can register only the languages the site needs
Renderer controlThe highlighter owned the final HTMLThe React component owns layout, copying, escaping, and controls

There is a measurable difference too. In TanStack’s comparison across 334 documentation fixtures, warmed highlighting took 4.6 ms with TanStack Highlight and 182 ms with Shiki 4.3.1. The generated HTML was 365 KiB versus 1,257 KiB. TanStack also reported no separate initialization or language-loading step in that comparison.

These are not measurements from this site, and the work is not perfectly equivalent. Shiki provides deeper grammar accuracy. Still, the numbers explain why the smaller output and synchronous path are useful for documentation. The bundle can become smaller again when I replace allLanguages with explicit imports for the languages this site actually uses.

This is also the main point from TanStack’s comparison: Shiki is still the stronger choice for VS Code-level accuracy and TextMate grammar compatibility. TanStack Highlight is a better fit for documentation where the languages are known and SSR, smaller output, and simple styling matter more.

Features that made the migration worthwhile

Themes are just CSS

Shiki usually puts colour values into the generated HTML. TanStack Highlight uses semantic classes such as th-keyword, th-string, and th-comment. CSS decides how they look.

This means the same markup works in both light and dark mode. Changing the theme does not require highlighting every code block again:

ts
const css = createThemeCss({
  light: githubLightTheme,
  dark: githubDarkTheme,
  darkSelector: '.dark',
})

I can also style the classes directly:

css
.th-keyword { color: var(--color-primary); }
.th-string { color: var(--color-success); }
.th-comment { color: var(--color-muted-foreground); font-style: italic; }

This makes the code blocks feel like part of the site instead of a separate editor theme.

Custom themes are easy to adjust

Themes are small semantic colour maps. I can start with an existing theme and change the few colours that matter to me:

ts
const onemanTheme = {
  ...githubDarkTheme,
  name: 'oneman-dark',
  background: '#101820',
  foreground: '#d9e2ec',
  tokens: {
    ...githubDarkTheme.tokens,
    keyword: '#7dd3fc',
    string: '#86efac',
    comment: '#94a3b8',
  },
}

The token names stay the same, so changing the visual style does not affect the parser or the React component.

Decorations cover the common custom cases

My old Shiki transformer changed the output one line at a time. TanStack Highlight has decorations for both lines and character ranges:

ts
const decorations = [
  { lines: [4, 6], className: 'normal' },
  { range: [12, 20], className: 'focus', data: { label: 'important' } },
]

Lines are one-based. Character ranges are zero-based and end-exclusive. This covers highlighted lines, diff lines, search matches, annotations, and small inline emphasis without writing another transformer.

I own the renderer

highlight() returns HTML when I need it. For this site, tokenize() and renderTokens() are more useful because React stays in control of the output.

That gives me control over escaping, line wrappers, diff signs, copy behaviour, and data attributes. TanStack Highlight only supplies the token information.

Embedded languages are explicit

HTML can pass its <script> and <style> bodies to the JavaScript and CSS tokenizers, but those languages need to be registered too:

ts
const highlighter = createHighlighter({
  languages: [html, css, js],
})

This is a useful rule to remember: register the languages that the outer language depends on. The same idea applies to Markdown fences.

Small custom languages are possible

For a small project-specific format, I do not need to build a full grammar. A custom language can return ranges with existing classes such as keyword, string, or property.

That could be useful for task lists, configuration files, command output, or a small DSL. I would rather leave uncertain text unclassified than give it a confidently wrong colour.

Tips I want to keep

  • Keep one highlighter at module scope. Recreating it for every block defeats the simple synchronous setup.
  • Register only the languages the application needs. This project currently uses allLanguages for convenience; selective imports are the next easy bundle improvement.
  • Register delegated languages explicitly. HTML without js and css will leave embedded bodies uncoloured.
  • Normalize aliases before rendering. bash, sh, and zsh can resolve to the registered shell language, while unknown languages should fall back to plaintext.
  • Keep line decorations in application CSS. The highlighter adds structure; spacing, borders, line numbers, diff colours, and copy controls belong to the code block component.
  • Use renderCodeBlockData when a component needs highlighted HTML, the normalized language, copy text, and tokens together.
  • Only inject HTML returned directly by TanStack Highlight. It escapes code text, but it is not a sanitizer for unrelated HTML in a Markdown document.
  • Keep the source text as the source of truth. Token values should reconstruct the original code, which keeps copying and testing predictable.

The tradeoff

TanStack Highlight is not a general replacement for Shiki. Shiki is still the better choice when I need exact VS Code tokenisation, TextMate grammar compatibility, semantic editor tokens, or hundreds of languages.

For this site, the languages are known and the output is documentation. I care more about predictable SSR, a simple rendering path, CSS-controlled themes, and small custom decorations. TanStack Highlight fits that shape better, and the migration removed more complexity than it added.

Found this useful? Share!