Every pandas column has one type for all its rows. One stray piece of text makes the whole column text, called object, and then the math stops working. This page loads seven employee expense claims with the usual mess in the amount column, shows what pd.to_numeric(..., errors="coerce") fixes and what it quietly destroys, and cleans the column properly. Homework 7 hands you a bills file with the same problem.
| claim_id | employee | category | amount |
|---|---|---|---|
| E-101 | Ana | Travel | 1200.50 |
| E-102 | Ben | Meals | $45.10 |
| E-103 | Cy | Travel | N/A |
| E-104 | Di | Lodging | 310 |
| E-105 | Ed | Meals | (blank) |
| E-106 | Flo | Travel | 1,050.00 |
| E-107 | Gus | Lodging | 289.99 |
errors="coerce" turns anything it cannot read into
NaN, a missing value, without telling you. Two of these claims are real money that coerce will throw
away. Finding them before they vanish is the skill.Run the cells top to bottom. Each one uses names from the cells above it, so skipping one gives a
NameError further down. Edit any cell and run it again.
Two functions for any messy amount column. Inside these boxes claims is the table above and pd is imported. The tests also run them on a different column.