Compare Two Lists in Python: Non-Matches, Duplicates and Order

To compare two lists in Python and return the non-matching elements, the shortest answer is the symmetric difference of two sets:

list(set(list_a) ^ set(list_b))       # items in one list but not the other

[x for x in list_a if x not in list_b]  # only the ones missing from b

Pick the second form when you need to know which list an item came from. Pick the first when you just want everything that differs.

One warning up front: sets discard duplicates, so both lines give wrong answers when your lists have repeats. There’s a section on that below.

All output on this page comes from real runs on Python 3.12.5.

Compare two lists and return non-matching elements

This is the question most people arrive with, so here it is both ways:

list_a = ["apple", "banana", "cherry", "date"]
list_b = ["banana", "date", "elderberry"]

# everything that is in one list but not the other
non_matches = list(set(list_a) ^ set(list_b))
print("non-matching:", sorted(non_matches))

# or keep the two directions separate, which is usually more useful
only_in_a = [x for x in list_a if x not in list_b]
only_in_b = [x for x in list_b if x not in list_a]

print("only in a   :", only_in_a)
print("only in b   :", only_in_b)

Output:

non-matching: ['apple', 'cherry', 'elderberry']
only in a   : ['apple', 'cherry']
only in b   : ['elderberry']
Command Prompt showing Python returning the non-matching elements of two lists using symmetric difference and list comprehensions
The symmetric difference merges both directions. The comprehensions keep them apart.

^ is the symmetric difference operator. set(a) ^ set(b) means everything in exactly one of the two sets.

Sets have no order, so wrap the result in sorted() if you need predictable output. Otherwise it may print differently between runs.

Comparing lists that contain duplicates

Here is where the set approach quietly breaks. Duplicates are thrown away before the comparison happens:

from collections import Counter

a = [1, 2, 2, 3, 4]
b = [2, 3, 3, 5]

print("a =", a)
print("b =", b)

# sets throw duplicates away
print("\nset(a) ^ set(b) :", set(a) ^ set(b), "  <- the second 2 and 3 vanished")

# Counter keeps the counts
count_a, count_b = Counter(a), Counter(b)
print("\nCounter a - b   :", list((count_a - count_b).elements()))
print("Counter b - a   :", list((count_b - count_a).elements()))
print("both directions :", list(((count_a - count_b) + (count_b - count_a)).elements()))

Output:

a = [1, 2, 2, 3, 4]
b = [2, 3, 3, 5]

set(a) ^ set(b) : {1, 4, 5}   <- the second 2 and 3 vanished

Counter a - b   : [1, 2, 4]
Counter b - a   : [3, 5]
both directions : [1, 2, 4, 3, 5]
Command Prompt comparing set symmetric difference against collections Counter on two Python lists containing duplicates
The set drops the extra 2 and 3. Counter keeps them.

Counter treats a list as a bag of counted items. Subtracting one from another leaves the surplus, which is what you usually want.

Use Counter whenever the number of times something appears carries meaning, such as stock levels, votes or log entries.

How to check if two lists are equal in Python

== compares position by position, so order matters. That’s often not what people mean by “equal”:

from collections import Counter

x = [1, 2, 3]
y = [3, 2, 1]

print("x == y                       :", x == y, "  (order matters)")
print("sorted(x) == sorted(y)       :", sorted(x) == sorted(y))

print()
# the trap: sets say two lists match when they do not
p = [1, 1, 2]
q = [1, 2]
print("p =", p, " q =", q)
print("set(p) == set(q)             :", set(p) == set(q), "  <- WRONG, p has two 1s")
print("Counter(p) == Counter(q)     :", Counter(p) == Counter(q), "  <- correct")

Output:

x == y                       : False   (order matters)
sorted(x) == sorted(y)       : True

p = [1, 1, 2]  q = [1, 2]
set(p) == set(q)             : True   <- WRONG, p has two 1s
Counter(p) == Counter(q)     : False   <- correct
Command Prompt showing that a Python set comparison wrongly reports two lists as equal when one contains duplicates
set(p) == set(q) says these match. They don’t.
You wantUse
Same items, same ordera == b
Same items, any ordersorted(a) == sorted(b)
Same items and same countsCounter(a) == Counter(b)
Same unique items onlyset(a) == set(b)

sorted() needs the items to be comparable, so it fails on mixed types. Counter only needs them hashable.

Finding common elements and differences

The four set operations cover almost every list comparison you’ll write:

a = ["red", "green", "blue", "yellow"]
b = ["blue", "yellow", "purple"]

print("in both      :", sorted(set(a) & set(b)))
print("in a only    :", sorted(set(a) - set(b)))
print("in b only    :", sorted(set(b) - set(a)))
print("in either    :", sorted(set(a) | set(b)))

print()
# keeping the original order of list a
print("common, a's order:", [x for x in a if x in set(b)])

Output:

in both      : ['blue', 'yellow']
in a only    : ['green', 'red']
in b only    : ['purple']
in either    : ['blue', 'green', 'purple', 'red', 'yellow']

common, a's order: ['blue', 'yellow']
  • & intersection — in both lists.
  • - difference — in the first but not the second.
  • | union — in either list, with duplicates removed.
  • ^ symmetric difference — in one list but not both.

The last line of that example is worth copying. Filtering the original list against a set keeps your ordering while still getting the fast lookup, which matters as soon as you care about sort order.

The fastest way to compare two large lists in Python

x not in some_list scans the whole list every time. x not in some_set is a single hash lookup:

import random
import timeit

big_a = list(range(20_000))
big_b = list(range(10_000, 30_000))
random.shuffle(big_b)

slow = timeit.timeit(lambda: [x for x in big_a if x not in big_b], number=1)

lookup = set(big_b)                      # convert ONCE, outside the loop
fast = timeit.timeit(lambda: [x for x in big_a if x not in lookup], number=1)

pure = timeit.timeit(lambda: set(big_a) - set(big_b), number=1)

print(f"20,000 items against 20,000 items\n")
print(f"  x not in list   {slow * 1000:>9.1f} ms")
print(f"  x not in set    {fast * 1000:>9.1f} ms   {slow / fast:>7,.0f}x faster")
print(f"  set difference  {pure * 1000:>9.1f} ms   {slow / pure:>7,.0f}x faster")

Output:

20,000 items against 20,000 items

  x not in list      1691.8 ms
  x not in set          0.4 ms     4,212x faster
  set difference        2.8 ms       607x faster
Command Prompt benchmark showing a Python set lookup running thousands of times faster than a list lookup when comparing two lists
Same answer, same loop. Only the container changed.

With 20,000 items in each list, the list version takes well over a second. The set version finishes in under a millisecond.

The reason is the shape of the work. Scanning a list is proportional to its length, so comparing two lists of n items does roughly n squared checks.

Convert the list you’re searching once, before the loop. Calling set() inside the comprehension rebuilds it on every iteration and throws the advantage away.

Comparing two lists element by element

When position matters, zip walks both lists together:

scores_before = [70, 82, 65, 90]
scores_after = [75, 82, 60, 95]
names = ["Ana", "Ben", "Cleo", "Dev"]

# zip walks both lists together, position by position
for name, before, after in zip(names, scores_before, scores_after):
    change = after - before
    mark = "same" if change == 0 else f"{change:+d}"
    print(f"{name:<6} {before:>3} -> {after:>3}  {mark}")

print()
# which positions differ at all
differing = [i for i, (x, y) in enumerate(zip(scores_before, scores_after)) if x != y]
print("positions that changed:", differing)

Output:

Ana     70 ->  75  +5
Ben     82 ->  82  same
Cleo    65 ->  60  -5
Dev     90 ->  95  +5

positions that changed: [0, 2, 3]

zip stops at the shorter list. If the lengths might differ and that matters, use itertools.zip_longest or check the lengths first.

enumerate alongside zip gives you the index of each mismatch, which is what you want for comparing strings character by character too.

Comparing lists of dictionaries

Sets need hashable items, and dictionaries aren’t hashable. This raises rather than returning a wrong answer:

# dictionaries cannot go in a set, so set() raises on a list of them
users_a = [{"id": 1, "name": "Ana"}, {"id": 2, "name": "Ben"}]
users_b = [{"id": 2, "name": "Ben"}, {"id": 3, "name": "Cleo"}]

try:
    set(users_a) - set(users_b)
except TypeError as err:
    print("set() on dicts ->", err)

print()
# compare on a key instead
ids_b = {u["id"] for u in users_b}
only_in_a = [u for u in users_a if u["id"] not in ids_b]
print("only in a:", only_in_a)

# or make each item hashable by freezing it
frozen_b = {tuple(sorted(u.items())) for u in users_b}
missing = [u for u in users_a if tuple(sorted(u.items())) not in frozen_b]
print("by whole record:", missing)

Output:

set() on dicts -> unhashable type: 'dict'

only in a: [{'id': 1, 'name': 'Ana'}]
by whole record: [{'id': 1, 'name': 'Ana'}]
Command Prompt showing a TypeError when building a set from Python dictionaries and two working alternatives
unhashable type: 'dict', then two ways around it.

Comparing on a single key is the usual answer, and it’s fast because the key set is built once.

Freezing each record into a sorted tuple compares whole dictionaries instead. It’s slower, but it catches changed values rather than just missing records.

Which Python list comparison method should you use?

GoalMethodKeeps duplicates?
Non-matching itemsset(a) ^ set(b)No
Non-matching, with duplicatesCounter subtractionYes
Items only in aset(a) - set(b)No
Common itemsset(a) & set(b)No
Equal ignoring orderCounter(a) == Counter(b)Yes
Position by positionzip(a, b)Yes
Lists of dictionariesCompare on a keyYes

Related list and comparison guides:

Frequently asked questions

How do I compare two lists in Python and return the non-matches?

Use list(set(a) ^ set(b)) for everything that differs, or two list comprehensions if you need to know which list each item came from. The set operators are covered in the Python set type reference.

How do I check if two lists are equal in Python?

a == b compares them in order. For order-insensitive equality use sorted(a) == sorted(b), or Counter(a) == Counter(b) when duplicates matter.

Why does my list comparison ignore duplicates?

Because converting to a set removes them. set([1, 1, 2]) == set([1, 2]) is True. Use collections.Counter instead.

How do I find elements in one list but not the other?

set(a) - set(b) gives items in a only. Swap the operands for the other direction.

What is the fastest way to compare two large lists?

Convert the list you are searching into a set once, before the loop. On 20,000 items that is over a thousand times faster than scanning the list.

How do I compare two lists of dictionaries?

Dictionaries are unhashable, so set() raises a TypeError. Compare on a key, or freeze each record with tuple(sorted(d.items())).

How do I compare two lists position by position?

Use zip(a, b), optionally with enumerate to get the index of each difference.