RMX: Meet Rio Terminal Multiplexer Protocol

No AI was used to write this article, therefore might have some English mistakes.

As most of developers knows, a terminal emulators are heavily based on text and a lot happens under the hood. If you have interest in a deep dive on the subject btw: I did a conference talk about terminals a while ago called “Terminals are just text? Think again” a while ago (there’s a lot of “like” repetition in my oratory, sorry I was nervous).

When you run ls, your shell writes some bytes to a thing called a pseudoterminal (pty for short). Think of a pty as a pipe with two ends: your application (often a shell) writes into one end, and your terminal reads out of the other 🤝.

Most of those bytes are plain text. But some of them are instructions, and those instructions are where all the interesting behavior lives. The instructions are called escape sequences, and they look like this when you print them raw:

\033[1;32mtests passed\033[0m

\033 is the escape character, it actually expect an operating system command (there’s other kinds of escape sequences also). The terminal reads that and understands:

  1. switch to bold green
  2. print “tests passed”
  3. then reset to normal

This is of course a simple overview of how it works. That’s basically how colors work, how cursor moves, how the screen clears, how images get displayed, etc. Everything the app wants the terminal to do is part of the stream of bytes as the text.

Then you have terminal multiplexers (like tmux), that follow the same principle.

If you see anyone using tmux, it’s likely that you will see someone using two or even more shells side by side in one window. Your terminal is basically drawing one grid of text, so how does tmux give you two?

All popular the terminal multiplexers sits in the middle.

  • To your shell, tmux pretends to be a terminal: it hands the shell a pty, reads the bytes, and interprets all those escape sequences itself, keeping its own in-memory grid of characters for each pane.
  • To your real terminal, tmux pretends to be an ordinary app: it takes its internal grids, composes them into one big picture (with the divider lines and the status bar), and re-emits new escape sequences describing that combined picture.

Here is the sequence, concretely, for a single colored word inside a pane:

  1. Your program writes \033[1;32mtests passed\033[0m to its pty.
  2. tmux parses it and records: “row 3, columns 0-7, text ‘tests passed’, bold, green” in its own grid.
  3. tmux decides what your real screen should look like, and writes fresh escape sequences: move cursor here, set bold green, print these characters.
  4. Your terminal parses those and finally draws pixels.
with tmux your app writes bytes 1 parse tmux grid copy #1 re-emits new escapes terminal grid copy #2 2nd parse pixels anything tmux cannot model is dropped here with rmx your app writes bytes same bytes, tagged with a buffer name nothing parses them on the way terminal grid only copy 1 parse pixels the pane is a real terminal, so images and new protocols survive
The same colored word, both ways. tmux parses it into its own grid and re-emits fresh escape sequences, so the terminal parses it a second time and anything tmux does not model is lost in between. With rmx the bytes are tagged with a buffer name and parsed once, by the terminal that draws them.