File Encoding — A Short Primer

This page explains one idea you meet in Lab 2: a text file can be perfectly good data and still refuse to open correctly, because of its encoding. It is a five-minute read.

What an encoding is

A text file is not really stored as letters. It is stored as bytes, meaning numbers. An encoding is the rulebook that maps those bytes back to characters you can read. Open a file with the wrong rulebook and the text comes out wrong: garbled, spaced out with gaps, or it fails to open at all. The data was fine; the rulebook was wrong.

The encodings you will meet

They are not interchangeable. A file written as UTF-16 has to be read as UTF-16.

Try it: type something, see the bytes

Type anything. Emoji and accented letters are where this gets interesting, because they are the characters that are not one byte each.

Try it: read the same bytes with the wrong rulebook

This is the failure you are here to recognize. The bytes never change. Only the rulebook used to read them changes.

The BOM — why the first line can look strange

Some files begin with a few invisible bytes called a byte-order mark (BOM) that announce the encoding. If your editor reads the file with the wrong rulebook, that BOM is what shows up as odd symbols at the very start of the first line. That is not corruption. It is the file telling you what it is.

Why it breaks an import

In Lab 2, JEA Detail.txt is UTF-16. A tool that assumes UTF-8 hits the very first bytes, cannot make sense of them, and stops before it has read a single row of data. Read the same file as a wrong single-byte encoding instead and every character comes back with a gap between it (C o m p a n y). The file is fine. The encoding was guessed wrong.

Characters you cannot see

Encoding is not the only thing in a file that is invisible. Tabs, line breaks and a handful of control characters are real characters that take up real bytes, and they are the reason a value can look blank while holding something. Lab 2 has one of these hiding in the ID column, and it is what makes a user count come out wrong.

Try it: make the invisible visible

Type or pick a value. Anything you cannot normally see is shown as a highlighted marker, with its byte underneath. Use \n, \t and \r to write them by hand.

Why this one costs you points. Try the Lab 2 ID bug preset. The value reads as BeanCounter25, and to a person it is BeanCounter25, but the stored value starts with a line break. All 117 of that user's rows carry it, so the count beside them is right. What breaks is the label: a pivot table shows the group with a blank or mangled name, and you cannot tell whose count it is. Excel's CLEAN removes the character and TRIM does not, because TRIM only removes spaces.

How to handle it

  1. Look first. Open the file in a text editor, which tells you the encoding. VS Code shows it in the blue bar at the bottom-right; Notepad shows it in the status bar at the bottom.
  2. In a Python script, state the encoding instead of letting the tool guess:
    import pandas as pd
    df = pd.read_csv("JEA Detail.txt", sep="\t", encoding="utf-16")
    If utf-16 does not work, try utf-16-le, utf-8, or latin-1 until the data reads cleanly.
  3. In Excel Power Query, the import preview has a File Origin box, which is the encoding. Power Query usually detects it for you automatically.
The Lab 2 point. Power Query detects the encoding for you. A Python script usually does not, because it does only what you tell it. So when you direct an AI to write a script, the encoding is a requirement your specification has to name out loud. That is the whole lesson in one detail.

Related: pandas primer · Lab 2 Output Validator · back to the Lab 2 page