
Why Your Spreadsheet Mangles José's Name
transcript
show notes
Ever opened a UTF-8 CSV in Excel and watched José turn into "José"? That garbled text is called mojibake, and it's the visible symptom of a character set mismatch hiding beneath every keyboard layout, font, and file format you touch. This episode unpacks the three layers people constantly collapse into one — code points, encodings, and graphemes — and shows why the encoding layer being wrong poisons everything above it. From ASCII's 1963 origins and UTF-8's genius backward compatibility, to the byte-length trap, the NFC/NFD normalization split that once made files vanish on macOS, and MySQL's years-long inability to store emoji, we trace how these failures reach real teams. The anchor is a story: a distributed operation preparing a several-hundred-person event across Israel, the UK, and New York, where a CRM export standardized on UTF-8 mangled the names of people about to receive an email. If you work with international data or collaborators, this is the layer you've been ignoring.
Episode #874473 — open it directly at myweirdprompts.com/874473





