On this page
You tried to read a text file and Python stopped with one of these:
UnicodeDecodeError: 'utf-8' codec can't decode byte 0xe9 in position 3: invalid continuation byteUnicodeDecodeError: 'utf-8' codec can't decode byte 0x80 in position 7: invalid start byteThe message sounds like a Python problem. It is a file problem: the file was not saved as UTF-8, and Python's default assumes it was. The fix is to tell Python the encoding the file really uses.
What the error means#
A text file is just bytes. An encoding is the agreement about which bytes mean which characters. UTF-8 is the most common one, and it is what open() assumed on the machine we tested. If the bytes follow a different agreement, UTF-8 decoding fails as soon as it meets a byte it cannot interpret.
The numbers in the message are useful. The byte is the one that broke decoding, and "position" is its offset in the file. We confirmed that offset on a test file:
first non-ascii byte index: 3 -> b'caf\xe9, 'The word café ended with the byte 0xe9, which is the letter é in Latin-1 but is not valid UTF-8. That one byte is your clue.
Step 1: Identify the encoding#
On macOS and Linux, file -I gives a first guess. We ran it on five files, each saved in a different encoding:
file -I latin1.txt cp1252.txt utf16.txt utf8bom.txt utf8.txtlatin1.txt: text/plain; charset=iso-8859-1
cp1252.txt: text/plain; charset=unknown-8bit
utf16.txt: text/plain; charset=utf-16le
utf8bom.txt: text/plain; charset=utf-8
utf8.txt: text/plain; charset=utf-8Four of those five were right, and one, cp1252.txt, came back as unknown-8bit. A tool can only guess. Use the clues in your error:
| What you see | Likely encoding |
|---|---|
A byte like 0xe9 in the middle of Western European text |
latin-1 or cp1252 |
0x80 to 0x9f bytes, curly quotes, a euro sign, files from Windows |
cp1252 |
Error at position 0, byte 0xff or 0xfe |
utf-16 |
| A strange invisible character at the start of the first line | utf-8-sig |
Step 2: Read it with the right encoding#
Pass the encoding to open(). Each of our test files then read correctly:
open('latin1.txt', encoding='latin-1').read() # 'café, naïve, Zoë\n'
open('cp1252.txt', encoding='cp1252').read() # 'price: €5 — “quoted”\n'
open('utf16.txt', encoding='utf-16').read() # 'hello utf-16\n'The same applies wherever text is read. csv.reader and bytes.decode() raised the same error on the same file, and the same encoding= argument fixes them. Libraries that read tabular data accept it too, as a keyword argument.
The invisible character: a byte order mark#
Some editors, mostly on Windows, put three bytes at the start of a UTF-8 file as a marker. Reading it as plain UTF-8 does not fail, which makes it sneakier. The marker comes through as a character you cannot see:
open('utf8bom.txt', encoding='utf-8').read()[:6]'name,'That ends up glued to your first column name, so row['name'] fails with a confusing KeyError. Use the utf-8-sig encoding, which strips it:
open('utf8bom.txt', encoding='utf-8-sig').read()[:6] # 'name,c'The silent trap: guessing latin-1 for a UTF-8 file#
latin-1 can decode any byte, so it never raises an error. That makes it tempting as a "just make it work" setting, and dangerous. Read a genuine UTF-8 file as Latin-1 and you get garbage with no warning:
open('utf8.txt', encoding='latin-1').read()'café â\x9c\x93\n'If your text suddenly shows é where an é should be, you decoded UTF-8 bytes with the wrong encoding. An error is better than this, because an error cannot corrupt your data without telling you.
Last resort: errors='replace' or 'ignore'#
If the data is only roughly text and some loss is acceptable, you can tell Python to continue past bad bytes:
open('latin1.txt', encoding='utf-8', errors='replace').read() # 'caf�, na�ve, Zo�\n'
open('latin1.txt', encoding='utf-8', errors='ignore').read() # 'caf, nave, Zo\n'Look at what happened to the words. replace swaps each bad byte for a replacement character, and ignore deletes it, turning café into caf. Both permanently damage the text. Use them for log files you only search, never for data you will write back out.
Find every bad byte in a big file#
For a large file, you may want to see all the offending spots rather than just the first:
data = open('latin1.txt', 'rb').read()
bad, pos = [], 0
while pos < len(data):
try:
data[pos:].decode('utf-8')
break
except UnicodeDecodeError as e:
bad.append((pos + e.start, data[pos + e.start]))
pos += e.end
print([(o, hex(b)) for o, b in bad])[(3, '0xe9'), (8, '0xef'), (15, '0xeb')]Those three bytes are the é, ï and ë in café, naïve, Zoë. If every bad byte is a single high byte inside otherwise ordinary text, the file is almost certainly Latin-1 or cp1252.
Stop it happening again#
- Always pass
encoding=explicitly when you open a text file. Do not rely on the platform default, which differs between machines. - Write files as UTF-8 unless you have a reason not to, and say so with
open(path, 'w', encoding='utf-8'). - When you receive a file from someone else, ask which program and system produced it.
Related reading: common character encoding issues, and writing a list to a file in Python, where choosing the encoding on the way out prevents this on the way in.