To read a binary file into a byte array in Python, use bytearray(Path("file.bin").read_bytes()), or open the file in "rb" mode and wrap f.read() in bytearray(). read() gives you immutable bytes; bytearray is the mutable version you can change and write back. This guide also covers reading byte by byte, reading large files in chunks, readinto(), memoryview, parsing a header with struct, and loading bytes into NumPy.
Every example was run with Python 3.12.5 and NumPy 2.5.3; the output is copied from the terminal. For the full reference, see open() in the Python documentation.
The sample file used in this guide
Run this once to create sample.bin, a 29-byte file with a small header, some text and a few raw bytes. Each example below starts from this file:
import struct
from pathlib import Path
# a small binary file: 10-byte header, some text, then 6 raw bytes
header = struct.pack("<4sHI", b"PGDT", 1, 42) # magic, version, record count
data = header + b"Hello, bytes!" + bytes(range(250, 256))
Path("sample.bin").write_bytes(data)
print(len(data), "bytes written to sample.bin")
Output:
29 bytes written to sample.bin
Quick answer: read the whole file into a bytearray
from pathlib import Path
data = bytearray(Path("sample.bin").read_bytes())
print(type(data).__name__, len(data), "bytes")
print(data[:4])
Output:
bytearray 29 bytes
bytearray(b'PGDT')
bytearray.Path.read_bytes() opens the file, reads it and closes it in one call. Wrapping the result in bytearray() gives you a mutable copy.
With open() and “rb” mode
The classic way works in every Python 3 version. The b in "rb" is essential: it returns raw bytes instead of trying to decode text:
with open("sample.bin", "rb") as f: # "rb" = read, binary
content = f.read() # bytes (immutable)
buffer = bytearray(content) # bytearray (mutable copy)
print(type(content).__name__, len(content))
print(type(buffer).__name__, len(buffer))
Output:
bytes 29
bytearray 29
bytes vs bytearray: which do you need?
Both hold the same numbers from 0 to 255. The difference is that bytearray can be changed:
from pathlib import Path
raw = Path("sample.bin").read_bytes()
buf = bytearray(raw)
buf[0] = ord("X") # a bytearray can be changed in place
print(buf[:4])
try:
raw[0] = ord("X") # bytes cannot
except TypeError as err:
print("TypeError:", err)
Output:
bytearray(b'XGDT')
TypeError: 'bytes' object does not support item assignment
bytes | bytearray | memoryview | |
|---|---|---|---|
| Returned by | f.read(), read_bytes() | bytearray(...), readinto() target | memoryview(obj) |
| Can be changed | No | Yes | Yes, if the object below it can |
| Slicing | Copies | Copies | No copy |
| Use it for | Reading, hashing, sending | Editing and writing back | Working on parts of big buffers |
Look inside the byte array
Indexing gives an int, slicing gives a new bytearray, and hex() and decode() turn bytes into readable text:
from pathlib import Path
data = bytearray(Path("sample.bin").read_bytes())
print(data[0]) # one item is an int from 0 to 255
print(list(data[:4])) # the first four bytes as numbers
print(data[:4] == b"PGDT") # compare with a bytes literal
print(data[10:23].decode("ascii"))
print(data[-6:].hex(" ")) # hex, one pair per byte
Output:
80
[80, 71, 68, 84]
True
Hello, bytes!
fa fb fc fd fe ff
A small hex dump, like the xxd tool, is handy for checking a file’s contents:
from pathlib import Path
data = Path("sample.bin").read_bytes()
for offset in range(0, len(data), 16):
chunk = data[offset:offset + 16]
hex_part = chunk.hex(" ").ljust(47)
text_part = "".join(chr(b) if 32 <= b < 127 else "." for b in chunk)
print(f"{offset:08x} {hex_part} {text_part}")
Output:
00000000 50 47 44 54 01 00 2a 00 00 00 48 65 6c 6c 6f 2c PGDT..*...Hello,
00000010 20 62 79 74 65 73 21 fa fb fc fd fe ff bytes!......
sample.bin, run in the Command Prompt.Change the bytes and save the file
from pathlib import Path
path = Path("sample.bin")
data = bytearray(path.read_bytes())
data[10:15] = b"HOWDY" # replace 5 bytes
data += b"\x00\x01" # append 2 bytes
path.write_bytes(data) # save it back
print(path.read_bytes()[10:23], len(path.read_bytes()), "bytes")
Output:
b'HOWDY, bytes!' 31 bytes
More on writing in write bytes to a file in Python.
Read a binary file byte by byte
f.read(1) returns one byte at a time and an empty b"" at the end of the file, which stops the while loop:
with open("sample.bin", "rb") as f:
count = 0
while byte := f.read(1): # b"" at the end of the file stops the loop
count += 1
if count <= 3:
print(byte, byte[0]) # a 1-byte bytes object and its int value
print(count, "bytes read one at a time")
Output:
b'P' 80
b'G' 71
b'D' 68
29 bytes read one at a time
This is slow for large files because of one call per byte; read chunks instead and loop over each chunk.
Read a large file in chunks
For files that don’t fit comfortably in memory, read fixed-size chunks. iter(callable, b"") keeps calling f.read() until it returns b"":
import hashlib
CHUNK = 8 # use 64 * 1024 or more for real files
sha = hashlib.sha256()
chunks = 0
with open("sample.bin", "rb") as f:
for chunk in iter(lambda: f.read(CHUNK), b""):
sha.update(chunk) # process each chunk, keep memory use small
chunks += 1
print(chunks, "chunks")
print(sha.hexdigest()[:16], "...")
Output:
4 chunks
3e070e6698a96043 ...
readinto(): fill an existing bytearray
readinto() writes directly into a buffer you created, instead of creating a new bytes object and copying it:
import os
size = os.path.getsize("sample.bin")
buffer = bytearray(size) # allocate once
with open("sample.bin", "rb") as f:
n = f.readinto(buffer) # fill it directly, no extra copy
print(n, "bytes read into a", type(buffer).__name__)
Output:
29 bytes read into a bytearray
memoryview: slice without copying
from pathlib import Path
data = bytearray(Path("sample.bin").read_bytes())
view = memoryview(data)
header = view[:10] # a slice of a memoryview does not copy the data
header[0:4] = b"ABCD" # writing through the view changes data too
print(data[:4])
print(view.nbytes, header.nbytes)
Output:
bytearray(b'ABCD')
29 10
Changing header changed data, because the view points at the same memory. That makes memoryview useful for parsing big files without copying slices.
Parse a binary header with struct
Most binary formats start with a fixed header. struct.unpack_from() turns those bytes back into Python values; <4sHI means little-endian, 4 raw bytes, an unsigned short and an unsigned int:
import struct
from pathlib import Path
data = Path("sample.bin").read_bytes()
magic, version, count = struct.unpack_from("<4sHI", data, 0)
print("magic:", magic)
print("version:", version)
print("record count:", count)
Output:
magic: b'PGDT'
version: 1
record count: 42
Load the bytes into a NumPy array
If you need numbers for calculations, np.fromfile() reads the file straight into an array. Change dtype to read 16- or 32-bit values:
import numpy as np
arr = np.fromfile("sample.bin", dtype=np.uint8) # every byte as a number
print(arr.dtype, arr.shape)
print(arr[:6])
print(arr[-6:])
Output:
uint8 (29,)
[80 71 68 84 1 0]
[250 251 252 253 254 255]
Common mistakes
Opening a binary file in text mode makes Python try to decode it as text, which fails as soon as it meets a byte that isn’t valid in the encoding:
try:
with open("sample.bin", "r", encoding="utf-8") as f: # text mode: wrong for binary data
f.read()
except UnicodeDecodeError as err:
print("UnicodeDecodeError:", err)
Output:
UnicodeDecodeError: 'utf-8' codec can't decode byte 0xfa in position 23: invalid start byte
A small wrapper can check the size first and give clear errors for missing or oversized files:
from pathlib import Path
def read_byte_array(path, max_bytes=100 * 1024 * 1024):
"""Read a whole file into a bytearray, refusing files larger than max_bytes."""
path = Path(path)
size = path.stat().st_size # raises FileNotFoundError if missing
if size > max_bytes:
raise ValueError(f"{path} is {size:,} bytes; read it in chunks instead")
return bytearray(path.read_bytes())
print(len(read_byte_array("sample.bin")))
for bad in ["missing.bin", "sample.bin"]:
try:
read_byte_array(bad, max_bytes=10)
except (FileNotFoundError, ValueError) as err:
print(type(err).__name__ + ":", err)
Output:
29
FileNotFoundError: [WinError 2] The system cannot find the file specified: 'missing.bin'
ValueError: sample.bin is 29 bytes; read it in chunks instead
The FileNotFoundError text above is from Windows; on macOS and Linux it reads [Errno 2] No such file or directory.
Which method should you use?
| Goal | Use |
|---|---|
| Whole file, change it | bytearray(Path(p).read_bytes()) |
| Whole file, read only | Path(p).read_bytes() |
| Very large file | iter(lambda: f.read(64 * 1024), b"") |
| No extra copies | f.readinto(bytearray(size)) + memoryview |
| Header or records | struct.unpack_from() |
| Numbers for maths | np.fromfile(p, dtype=...) |
Frequently asked questions
How do I read a binary file into a byte array in Python?
Use bytearray(Path("file.bin").read_bytes()), or with open("file.bin", "rb") as f: data = bytearray(f.read()).
What is the difference between bytes and bytearray?
Both store values from 0 to 255. bytes is immutable; bytearray can be changed in place, appended to and written back.
How do I read a file as bytes in Python?
Open it in binary mode: open(path, "rb").read() or Path(path).read_bytes(). Both return a bytes object.
How do I read a large binary file without running out of memory?
Read it in chunks with iter(lambda: f.read(64 * 1024), b"") and process each chunk, or use readinto() with a reusable buffer.
Why do I get UnicodeDecodeError when reading a binary file?
The file was opened in text mode ("r"). Use "rb" so Python returns raw bytes instead of decoding them.
Continue with these file-handling tutorials:
Bijay Kumar is a 13-time Microsoft MVP with more than 18 years in software development, and the founder of Python Guides and TSinfo Technologies. He started out building .NET and SharePoint solutions at HP, TCS and KPIT before moving into Python, machine learning and AI, and he also builds web apps with TypeScript and React. He writes the tutorials here himself, and every example is run before publishing so you see the real output. More about Bijay · Microsoft MVP profile · LinkedIn