This page explains one idea you meet in Lab 2: a text file can be perfectly good data and still refuse to open correctly, because of its encoding. It is a five-minute read.
A text file is not really stored as letters. It is stored as bytes, meaning numbers. An encoding is the rulebook that maps those bytes back to characters you can read. Open a file with the wrong rulebook and the text comes out wrong: garbled, spaced out with gaps, or it fails to open at all. The data was fine; the rulebook was wrong.
JEA Detail.txt in
Lab 2 uses.They are not interchangeable. A file written as UTF-16 has to be read as UTF-16.
Type anything. Emoji and accented letters are where this gets interesting, because they are the characters that are not one byte each.
This is the failure you are here to recognize. The bytes never change. Only the rulebook used to read them changes.
Some files begin with a few invisible bytes called a byte-order mark (BOM) that announce the encoding. If your editor reads the file with the wrong rulebook, that BOM is what shows up as odd symbols at the very start of the first line. That is not corruption. It is the file telling you what it is.
In Lab 2, JEA Detail.txt is UTF-16. A tool that assumes UTF-8 hits
the very first bytes, cannot make sense of them, and stops before it has read a
single row of data. Read the same file as a wrong single-byte encoding instead and
every character comes back with a gap between it (C o m p a n y). The
file is fine. The encoding was guessed wrong.
Encoding is not the only thing in a file that is invisible. Tabs, line breaks
and a handful of control characters are real characters that take up real bytes,
and they are the reason a value can look blank while holding something. Lab 2 has
one of these hiding in the ID column, and it is what makes a user
count come out wrong.
Type or pick a value. Anything you cannot normally see is shown
as a highlighted marker, with its byte underneath. Use \n,
\t and \r to write them by hand.
BeanCounter25, and to a person it is
BeanCounter25, but the stored value starts with a line break. All 117 of that
user's rows carry it, so the count beside them is right. What breaks is the
label: a pivot table shows the group with a blank or mangled name, and you
cannot tell whose count it is. Excel's CLEAN removes the
character and TRIM does not, because TRIM only
removes spaces.
import pandas as pd
df = pd.read_csv("JEA Detail.txt", sep="\t", encoding="utf-16")
If utf-16 does not work, try utf-16-le,
utf-8, or latin-1 until the data reads cleanly.Related: pandas primer · Lab 2 Output Validator · back to the Lab 2 page