PYTHON

Grouping and Aggregating Data with Python's collections.defaultdict

Leverage collections.defaultdict in Python to efficiently group items and aggregate values (e.g., sum, count) based on a key, simplifying complex data processing tasks.

from collections import defaultdict

# Example: Grouping sales data by region and summing total sales
sales_data = [
    {'region': 'North', 'product': 'A', 'sales': 100},
    {'region': 'South', 'product': 'B', 'sales': 150},
    {'region': 'North', 'product': 'C', 'sales': 200},
    {'region': 'East', 'product': 'A', 'sales': 50},
    {'region': 'South', 'product': 'D', 'sales': 75},
]

# Grouping and summing sales using defaultdict
regional_sales_total = defaultdict(int)
products_by_region = defaultdict(list)

for record in sales_data:
    region = record['region']
    sales = record['sales']
    product = record['product']

    regional_sales_total[region] += sales
    products_by_region[region].append(product)

print(f"Total sales by region: {dict(regional_sales_total)}")
print(f"Products by region: {dict(products_by_region)}")

# Another example: Grouping words by their first letter
words = ['apple', 'banana', 'grape', 'apricot', 'berry', 'guava']
words_by_initial = defaultdict(list)
for word in words:
    words_by_initial[word[0]].append(word)
print(f"Words grouped by initial: {dict(words_by_initial)}")
How it works: The `collections.defaultdict` is a specialized dictionary subclass that provides a default value for a key that hasn't been set. This significantly simplifies grouping and aggregation logic by eliminating the need to check if a key already exists before appending to a list or summing a value. When a new key is accessed, it automatically creates an entry with the default factory's return value (e.g., an empty list for `list` or `0` for `int`), making the code cleaner and less verbose.

Need help integrating this into your project?

Our team of expert developers can help you build your custom application from scratch.

Hire DigitalCodeLabs