How to Read a Binary File into a Byte Array in Python

To read a binary file into a byte array in Python, use bytearray(Path("file.bin").read_bytes()), or open the file in "rb" mode and wrap f.read() in bytearray(). read() gives you immutable bytes; bytearray is the mutable version you can change and write back. This guide also covers reading byte by byte, reading large files in chunks, readinto(), memoryview, parsing a header with struct, and loading bytes into NumPy.

Every example was run with Python 3.12.5 and NumPy 2.5.3; the output is copied from the terminal. For the full reference, see open() in the Python documentation.

The sample file used in this guide

Run this once to create sample.bin, a 29-byte file with a small header, some text and a few raw bytes. Each example below starts from this file:

import struct
from pathlib import Path

# a small binary file: 10-byte header, some text, then 6 raw bytes
header = struct.pack("<4sHI", b"PGDT", 1, 42)     # magic, version, record count
data = header + b"Hello, bytes!" + bytes(range(250, 256))

Path("sample.bin").write_bytes(data)
print(len(data), "bytes written to sample.bin")

Output:

29 bytes written to sample.bin
Command Prompt screenshot writing a 29-byte binary file with struct.pack and Path.write_bytes in Python
Creating the sample binary file used in this tutorial.

Quick answer: read the whole file into a bytearray

from pathlib import Path

data = bytearray(Path("sample.bin").read_bytes())

print(type(data).__name__, len(data), "bytes")
print(data[:4])

Output:

bytearray 29 bytes
bytearray(b'PGDT')
Command Prompt screenshot reading a binary file into a bytearray in Python with pathlib
Reading the file into a bytearray.

Path.read_bytes() opens the file, reads it and closes it in one call. Wrapping the result in bytearray() gives you a mutable copy.

With open() and “rb” mode

The classic way works in every Python 3 version. The b in "rb" is essential: it returns raw bytes instead of trying to decode text:

with open("sample.bin", "rb") as f:     # "rb" = read, binary
    content = f.read()                  # bytes (immutable)

buffer = bytearray(content)             # bytearray (mutable copy)

print(type(content).__name__, len(content))
print(type(buffer).__name__, len(buffer))

Output:

bytes 29
bytearray 29

bytes vs bytearray: which do you need?

Both hold the same numbers from 0 to 255. The difference is that bytearray can be changed:

from pathlib import Path

raw = Path("sample.bin").read_bytes()
buf = bytearray(raw)

buf[0] = ord("X")               # a bytearray can be changed in place
print(buf[:4])

try:
    raw[0] = ord("X")           # bytes cannot
except TypeError as err:
    print("TypeError:", err)

Output:

bytearray(b'XGDT')
TypeError: 'bytes' object does not support item assignment
bytesbytearraymemoryview
Returned byf.read(), read_bytes()bytearray(...), readinto() targetmemoryview(obj)
Can be changedNoYesYes, if the object below it can
SlicingCopiesCopiesNo copy
Use it forReading, hashing, sendingEditing and writing backWorking on parts of big buffers

Look inside the byte array

Indexing gives an int, slicing gives a new bytearray, and hex() and decode() turn bytes into readable text:

from pathlib import Path

data = bytearray(Path("sample.bin").read_bytes())

print(data[0])                  # one item is an int from 0 to 255
print(list(data[:4]))           # the first four bytes as numbers
print(data[:4] == b"PGDT")      # compare with a bytes literal
print(data[10:23].decode("ascii"))
print(data[-6:].hex(" "))       # hex, one pair per byte

Output:

80
[80, 71, 68, 84]
True
Hello, bytes!
fa fb fc fd fe ff

A small hex dump, like the xxd tool, is handy for checking a file’s contents:

from pathlib import Path

data = Path("sample.bin").read_bytes()

for offset in range(0, len(data), 16):
    chunk = data[offset:offset + 16]
    hex_part = chunk.hex(" ").ljust(47)
    text_part = "".join(chr(b) if 32 <= b < 127 else "." for b in chunk)
    print(f"{offset:08x}  {hex_part}  {text_part}")

Output:

00000000  50 47 44 54 01 00 2a 00 00 00 48 65 6c 6c 6f 2c  PGDT..*...Hello,
00000010  20 62 79 74 65 73 21 fa fb fc fd fe ff            bytes!......
Command Prompt screenshot of a Python hex dump of a binary file showing offsets, hex bytes and ASCII text
The hex dump of sample.bin, run in the Command Prompt.

Change the bytes and save the file

from pathlib import Path

path = Path("sample.bin")
data = bytearray(path.read_bytes())

data[10:15] = b"HOWDY"          # replace 5 bytes
data += b"\x00\x01"             # append 2 bytes

path.write_bytes(data)          # save it back
print(path.read_bytes()[10:23], len(path.read_bytes()), "bytes")

Output:

b'HOWDY, bytes!' 31 bytes

More on writing in write bytes to a file in Python.

Read a binary file byte by byte

f.read(1) returns one byte at a time and an empty b"" at the end of the file, which stops the while loop:

with open("sample.bin", "rb") as f:
    count = 0
    while byte := f.read(1):        # b"" at the end of the file stops the loop
        count += 1
        if count <= 3:
            print(byte, byte[0])    # a 1-byte bytes object and its int value
print(count, "bytes read one at a time")

Output:

b'P' 80
b'G' 71
b'D' 68
29 bytes read one at a time

This is slow for large files because of one call per byte; read chunks instead and loop over each chunk.

Read a large file in chunks

For files that don’t fit comfortably in memory, read fixed-size chunks. iter(callable, b"") keeps calling f.read() until it returns b"":

import hashlib

CHUNK = 8                                   # use 64 * 1024 or more for real files

sha = hashlib.sha256()
chunks = 0
with open("sample.bin", "rb") as f:
    for chunk in iter(lambda: f.read(CHUNK), b""):
        sha.update(chunk)                   # process each chunk, keep memory use small
        chunks += 1

print(chunks, "chunks")
print(sha.hexdigest()[:16], "...")

Output:

4 chunks
3e070e6698a96043 ...

readinto(): fill an existing bytearray

readinto() writes directly into a buffer you created, instead of creating a new bytes object and copying it:

import os

size = os.path.getsize("sample.bin")
buffer = bytearray(size)                    # allocate once

with open("sample.bin", "rb") as f:
    n = f.readinto(buffer)                  # fill it directly, no extra copy

print(n, "bytes read into a", type(buffer).__name__)

Output:

29 bytes read into a bytearray

memoryview: slice without copying

from pathlib import Path

data = bytearray(Path("sample.bin").read_bytes())
view = memoryview(data)

header = view[:10]                  # a slice of a memoryview does not copy the data
header[0:4] = b"ABCD"               # writing through the view changes data too

print(data[:4])
print(view.nbytes, header.nbytes)

Output:

bytearray(b'ABCD')
29 10

Changing header changed data, because the view points at the same memory. That makes memoryview useful for parsing big files without copying slices.

Parse a binary header with struct

Most binary formats start with a fixed header. struct.unpack_from() turns those bytes back into Python values; <4sHI means little-endian, 4 raw bytes, an unsigned short and an unsigned int:

import struct
from pathlib import Path

data = Path("sample.bin").read_bytes()

magic, version, count = struct.unpack_from("<4sHI", data, 0)
print("magic:", magic)
print("version:", version)
print("record count:", count)

Output:

magic: b'PGDT'
version: 1
record count: 42

Load the bytes into a NumPy array

If you need numbers for calculations, np.fromfile() reads the file straight into an array. Change dtype to read 16- or 32-bit values:

import numpy as np

arr = np.fromfile("sample.bin", dtype=np.uint8)   # every byte as a number
print(arr.dtype, arr.shape)
print(arr[:6])
print(arr[-6:])

Output:

uint8 (29,)
[80 71 68 84  1  0]
[250 251 252 253 254 255]

Common mistakes

Opening a binary file in text mode makes Python try to decode it as text, which fails as soon as it meets a byte that isn’t valid in the encoding:

try:
    with open("sample.bin", "r", encoding="utf-8") as f:    # text mode: wrong for binary data
        f.read()
except UnicodeDecodeError as err:
    print("UnicodeDecodeError:", err)

Output:

UnicodeDecodeError: 'utf-8' codec can't decode byte 0xfa in position 23: invalid start byte

A small wrapper can check the size first and give clear errors for missing or oversized files:

from pathlib import Path

def read_byte_array(path, max_bytes=100 * 1024 * 1024):
    """Read a whole file into a bytearray, refusing files larger than max_bytes."""
    path = Path(path)
    size = path.stat().st_size                  # raises FileNotFoundError if missing
    if size > max_bytes:
        raise ValueError(f"{path} is {size:,} bytes; read it in chunks instead")
    return bytearray(path.read_bytes())

print(len(read_byte_array("sample.bin")))

for bad in ["missing.bin", "sample.bin"]:
    try:
        read_byte_array(bad, max_bytes=10)
    except (FileNotFoundError, ValueError) as err:
        print(type(err).__name__ + ":", err)

Output:

29
FileNotFoundError: [WinError 2] The system cannot find the file specified: 'missing.bin'
ValueError: sample.bin is 29 bytes; read it in chunks instead

The FileNotFoundError text above is from Windows; on macOS and Linux it reads [Errno 2] No such file or directory.

Which method should you use?

GoalUse
Whole file, change itbytearray(Path(p).read_bytes())
Whole file, read onlyPath(p).read_bytes()
Very large fileiter(lambda: f.read(64 * 1024), b"")
No extra copiesf.readinto(bytearray(size)) + memoryview
Header or recordsstruct.unpack_from()
Numbers for mathsnp.fromfile(p, dtype=...)

Frequently asked questions

How do I read a binary file into a byte array in Python?

Use bytearray(Path("file.bin").read_bytes()), or with open("file.bin", "rb") as f: data = bytearray(f.read()).

What is the difference between bytes and bytearray?

Both store values from 0 to 255. bytes is immutable; bytearray can be changed in place, appended to and written back.

How do I read a file as bytes in Python?

Open it in binary mode: open(path, "rb").read() or Path(path).read_bytes(). Both return a bytes object.

How do I read a large binary file without running out of memory?

Read it in chunks with iter(lambda: f.read(64 * 1024), b"") and process each chunk, or use readinto() with a reusable buffer.

Why do I get UnicodeDecodeError when reading a binary file?

The file was opened in text mode ("r"). Use "rb" so Python returns raw bytes instead of decoding them.

Continue with these file-handling tutorials: