Partition data without creating hotspots
Choose partition keys and recognize skew, fan-out, and rebalancing costs.
- Compare range and hash partitioning and identify a hot partition.
Partitioning divides data across nodes so storage and work can grow horizontally. Range partitioning keeps nearby keys together and can support ordered scans, but writes to the newest range may concentrate on one node. Hash partitioning spreads keys more evenly, but makes range scans harder. A popular key can still overload one partition. Consistent hashing can reduce the amount of data moved when membership changes, but it does not remove skew by itself.
events = {"partition-a": 12, "partition-b": 11, "partition-c": 90}
hot = max(events, key=events.get)
print(f"hot partition: {hot} ({events[hot]} events)")hot partition: partition-c (90 events)
This Python example counts events per partition. Possible responses to skew include choosing a better key, salting very hot keys, splitting a tenant, or adding a specialized serving path; each affects query patterns and operations.
Key takeaways
Choose a partition key based on access patterns and key distribution.
Range and hash partitioning trade locality against spread.
Consistent hashing helps rebalance membership changes, not workload skew.
Lesson quiz
5 questions · pass with 4 correct · up to 50 XP
Passing this quiz completes the lesson and keeps your streak going. Questions you miss come back in review sessions later.
Questions about this lesson
Stuck? Ask. Figured something out? Share it. Explaining is one of the best ways to learn.
Loading posts…