PYTHON
Grouping and Aggregating Data with Python's collections.defaultdict
Leverage collections.defaultdict in Python to efficiently group items and aggregate values (e.g., sum, count) based on a key, simplifying complex data processing tasks.
from collections import defaultdict
# Example: Grouping sales data by region and summing total sales
sales_data = [
{'region': 'North', 'product': 'A', 'sales': 100},
{'region': 'South', 'product': 'B', 'sales': 150},
{'region': 'North', 'product': 'C', 'sales': 200},
{'region': 'East', 'product': 'A', 'sales': 50},
{'region': 'South', 'product': 'D', 'sales': 75},
]
# Grouping and summing sales using defaultdict
regional_sales_total = defaultdict(int)
products_by_region = defaultdict(list)
for record in sales_data:
region = record['region']
sales = record['sales']
product = record['product']
regional_sales_total[region] += sales
products_by_region[region].append(product)
print(f"Total sales by region: {dict(regional_sales_total)}")
print(f"Products by region: {dict(products_by_region)}")
# Another example: Grouping words by their first letter
words = ['apple', 'banana', 'grape', 'apricot', 'berry', 'guava']
words_by_initial = defaultdict(list)
for word in words:
words_by_initial[word[0]].append(word)
print(f"Words grouped by initial: {dict(words_by_initial)}")
How it works: The `collections.defaultdict` is a specialized dictionary subclass that provides a default value for a key that hasn't been set. This significantly simplifies grouping and aggregation logic by eliminating the need to check if a key already exists before appending to a list or summing a value. When a new key is accessed, it automatically creates an entry with the default factory's return value (e.g., an empty list for `list` or `0` for `int`), making the code cleaner and less verbose.