GIC Engineering Consultants
Home Articles Services Contact
Storage Architecture: When Bucket Strategy Goes Bad

Storage Architecture: When Bucket Strategy Goes Bad

By Marcus House, Splunk Enterprise Architect

"Why is our hot storage at 95% capacity?"

I looked at their Splunk environment. 30 TB of storage across the cluster. Plenty of space.

Except 90% of it was in cold buckets. Hot storage had 48 hours until complete failure.

Understanding Splunk's Bucket Strategy

Splunk data moves through a lifecycle:

The problem? Most teams set this up once during installation and never revisit it. Then data patterns change, volume grows, and suddenly hot storage fills up.

Real Case: The STIG Disaster

One of my early projects involved a Splunk environment where someone had implemented STIG requirements too aggressively. They set the minimum free disk space to 50%.

Splunk hit that threshold and stopped indexing—even though 12 TB of space was still available.

Users started calling: "Data isn't showing up in searches."

Security team panicked: "Are we under attack?"

Reality: Bad storage configuration, not a breach.

We overrode the STIG setting to something practical (15% minimum free space). Indexing resumed immediately.

The Three Configuration Mistakes

1. Not Using Separate Volumes for Hot/Warm vs. Cold

Many teams put all buckets on the same volume. This means:

Correct approach:

[volume:hot_warm]
path = /mnt/fast_ssd/splunk_hot_warm
maxVolumeDataSizeMB = 500000

[volume:cold]
path = /mnt/cheap_sata/splunk_cold
maxVolumeDataSizeMB = 10000000

2. Setting Max Bucket Size Too Small

Default: 750 MB per bucket (auto)

Some teams set this to 100 MB thinking "smaller buckets = better performance."

Wrong. This creates:

I recently worked on an environment with 450,000 buckets because maxDataSize was set to 50 MB. Searches were painfully slow. We increased to 1 GB and bucket count dropped to 22,000. Search performance improved 300%.

3. Not Planning for Growth

Today's 1 GB/day can become tomorrow's 100 GB/day. I've seen:

Always provision storage for 3x current daily volume. If you're ingesting 100 GB/day, plan for 300 GB/day capacity.

The Storage Formula

Here's how I calculate storage requirements:

Daily Ingest (GB/day) × Retention (days) × RF (replication factor) × 1.2 (buffer) = Total Storage Needed

Example:

200 × 90 × 3 × 1.2 = 64,800 GB = ~65 TB

The Monitoring You Need

Set up alerts for:

Hot/Warm Storage Capacity:

| rest splunk_server=local /services/server/status/partitions-space | eval pct_used=round((1-(free/capacity))*100, 1) | where pct_used > 80 | table splunk_server title capacity free pct_used

Cold Storage Capacity:

| dbinspect index=* | stats sum(sizeOnDiskMB) as size_mb by index state splunk_server | where state="cold" | eval size_gb=round(size_mb/1024, 2) | sort -size_gb

Bucket Aging Issues:

index=_internal source=*splunkd.log "Not rolling" OR "Not freezing" | stats count by host

I run these searches daily. When storage hits 80%, I have at least a week to add capacity before hitting critical levels.

The Index-Level Optimization

Not all indexes need the same retention. I use tiered retention:

[security_critical]
homePath = volume:hot_warm/security_critical/db
coldPath = volume:cold/security_critical/colddb
maxDataSizeMB = 1024
frozenTimePeriodInSecs = 31536000  # 365 days

[debug_logs]
homePath = volume:hot_warm/debug_logs/db
coldPath = volume:cold/debug_logs/colddb
maxDataSizeMB = 512
frozenTimePeriodInSecs = 604800  # 7 days

Security logs: 1 year retention

Debug logs: 7 days

This prevents debug noise from consuming storage needed for critical security data.

The Cleanup Script

For indexes with excessive retention, set the proper retention policy in indexes.conf:

[old_app_logs]
frozenTimePeriodInSecs = 7776000
# 90 days = 90 * 24 * 60 * 60 = 7,776,000 seconds
# Data older than this will automatically roll to frozen

To identify candidates for retention adjustment:

| dbinspect index=old_app_logs | eval age_days=round((now()-endEpoch)/86400, 0) | where age_days > 90 | stats sum(sizeOnDiskMB) as recoverable_mb count as bucket_count by splunk_server

The Takeaway

Storage problems in Splunk are never sudden. They're always predictable, always preventable, and always fixable—if you're monitoring proactively.

If you're not monitoring storage capacity daily, you will eventually have an incident. It's not "if," it's "when."

Have you dealt with storage disasters in Splunk? Share your story below.

I publish a new Splunk article every week. Follow me to catch the next one.

← Back to Articles