The essentials
Quick reference
One focused task per row. Jump to the related section for complete, working examples.
| Use | Syntax | Examples |
|---|---|---|
| Parse exact timestamps | df['time'] = pd.to_datetime(df['time'], format='%Y-%m-%d %H:%M', errors='raise') | View examples |
| Normalize to UTC | df['time'] = pd.to_datetime(df['time'], utc=True, errors='coerce') | View examples |
| Create a regular index | index = pd.date_range('2026-08-01', periods=24, freq='h', tz='UTC') | View examples |
| Set a sorted time index | series = df.set_index('time').sort_index() | View examples |
| Extract time components | df['weekday'] = df['time'].dt.day_name() | View examples |
| Localize wall times | localized = naive.dt.tz_localize('America/Sao_Paulo', ambiguous='raise', nonexistent='raise') | View examples |
| Convert time zones | local = utc.dt.tz_convert('America/Sao_Paulo') | View examples |
| Aggregate daily totals | daily = series.resample('D').sum(min_count=1) | View examples |
| Control interval edges | daily = series.resample('D', closed='right', label='right').sum() | View examples |
| Expose missing intervals | hourly = series.asfreq('h') | View examples |
| Interpolate by time | hourly = series.resample('h').asfreq().interpolate(method='time', limit_area='inside') | View examples |
| Calculate a rolling mean | result = series.rolling('3h', min_periods=2).mean() | View examples |
| Calculate an expanding total | result = series.expanding(min_periods=1).sum() | View examples |
| Calculate an EWM mean | result = series.ewm(span=3, adjust=False).mean() | View examples |
| Create a one-period lag | df['previous'] = df['value'].shift(1) | View examples |
| Calculate fractional change | df['change'] = df['value'].pct_change(fill_method=None) | View examples |
Time-series analysis depends on an explicit time contract: parse known formats, distinguish localization from conversion, sort chronological indexes, and define interval boundaries and window requirements before calculating aggregates or changes.
Step by step
Detailed examples
Parse timestamps into a chronological index
Use an exact format and errors='raise' when invalid input should stop ingestion. errors='coerce' supports data-quality reporting because failures become NaT. utc=True normalizes aware inputs to a common timeline. Most time-aware operations expect a DatetimeIndex or an explicit on column, and chronological sorting is required for reliable slicing and ordered calculations.
import pandas as pd
df = pd.DataFrame({
'time': ['2026-08-01 09:00', '2026-08-01 10:00', '2026-08-01 11:00'],
'value': [10, 14, 13],
})
df['time'] = pd.to_datetime(df['time'], format='%Y-%m-%d %H:%M', errors='raise')
series = df.set_index('time').sort_index()['value']
index = pd.date_range('2026-08-01', periods=3, freq='h', tz='UTC')
parsed_utc = pd.to_datetime(pd.Series(['2026-08-01T09:00:00-03:00']), utc=True, errors='coerce')
print(series.index.is_monotonic_increasing)
print(index.astype(str).tolist())
print(parsed_utc.astype(str).tolist()) True
['2026-08-01 00:00:00+00:00', '2026-08-01 01:00:00+00:00', '2026-08-01 02:00:00+00:00']
['2026-08-01 12:00:00+00:00']Distinguish localization from conversion
tz_localize assigns a zone to naive wall-clock values, while tz_convert changes the displayed zone of already-aware instants. Daylight-saving transitions can create ambiguous or nonexistent local times; choose explicit handling instead of silently guessing. Extract calendar components after converting to the reporting zone because weekdays and dates can differ across zones.
import pandas as pd
naive = pd.Series(pd.to_datetime(['2026-08-01 09:00', '2026-08-02 10:00']))
local = naive.dt.tz_localize(
'America/Sao_Paulo', ambiguous='raise', nonexistent='raise',
)
utc = local.dt.tz_convert('UTC')
reported = utc.dt.tz_convert('America/Sao_Paulo')
weekday = reported.dt.day_name()
print(utc.astype(str).tolist())
print(weekday.tolist()) ['2026-08-01 12:00:00+00:00', '2026-08-02 13:00:00+00:00']
['Saturday', 'Sunday']Define resampling bins and gap policy
resample groups observations into time bins and then requires an aggregation or fill operation. closed controls which boundary owns an observation and label controls the timestamp shown for the bin. asfreq only conforms data to a frequency, exposing gaps without aggregation. Interpolation invents estimates, so limit it to bounded internal gaps and use it only for measures where interpolation is defensible.
import pandas as pd
index = pd.to_datetime([
'2026-08-01 00:00', '2026-08-01 02:00', '2026-08-02 00:00',
])
series = pd.Series([10.0, 14.0, 20.0], index=index)
daily = series.resample('D').sum(min_count=1)
right_labeled = series.resample('D', closed='right', label='right').sum()
hourly_gaps = series.asfreq('h')
interpolated = series.resample('h').asfreq().interpolate(method='time', limit_area='inside')
print(daily.to_dict())
print(right_labeled.index.astype(str).tolist())
print(hourly_gaps.iloc[:3].tolist())
print(interpolated.iloc[:3].tolist()) {Timestamp('2026-08-01 00:00:00'): 24.0, Timestamp('2026-08-02 00:00:00'): 20.0}
['2026-08-01', '2026-08-02', '2026-08-03']
[10.0, nan, 14.0]
[10.0, 12.0, 14.0]Choose fixed, expanding, or weighted windows
A time-based rolling window covers elapsed time and therefore requires a monotonic datetime-like index, while an integer window covers a number of rows. min_periods controls when results become valid. expanding includes all observations so far, and ewm applies exponentially decreasing weights; select parameters from the analytical meaning rather than tuning until the result looks smooth.
import pandas as pd
series = pd.Series(
[10.0, 14.0, 13.0, 19.0],
index=pd.date_range('2026-08-01 09:00', periods=4, freq='h'),
)
rolling = series.rolling('3h', min_periods=2).mean()
expanding = series.expanding(min_periods=1).sum()
weighted = series.ewm(span=3, adjust=False).mean()
print(rolling.round(2).tolist())
print(expanding.tolist())
print(weighted.round(2).tolist()) [nan, 12.0, 12.33, 15.33]
[10.0, 24.0, 37.0, 56.0]
[10.0, 12.0, 12.5, 15.75]Align prior observations before measuring change
shift moves values by row position without changing the index, making the comparison alignment visible. pct_change returns fractional rather than percentage change, so multiply by 100 only for display as a percentage. Sort first, group by independent series when needed, and keep fill_method=None so missing observations do not silently become carried-forward values.
import pandas as pd
df = pd.DataFrame({
'time': pd.date_range('2026-08-01', periods=4, freq='D'),
'value': [100.0, 110.0, None, 121.0],
}).sort_values('time')
df['previous'] = df['value'].shift(1)
df['change'] = df['value'].pct_change(fill_method=None)
df['change_percent'] = df['change'].mul(100).round(2)
print(df[['value', 'previous', 'change_percent']].to_dict('records')) [{'value': 100.0, 'previous': nan, 'change_percent': nan}, {'value': 110.0, 'previous': 100.0, 'change_percent': 10.0}, {'value': nan, 'previous': 110.0, 'change_percent': nan}, {'value': 121.0, 'previous': nan, 'change_percent': nan}]Local code tester
Resample and smooth hourly observations
Edit the time-series values and compare daily aggregation with a rolling mean.
Press Run to load Python locally.
Sources and further reading
References
Authoritative documentation used to verify and expand this cheat sheet.
Help us improve
Found a typo or missing example?
Tell us what would make this cheat sheet clearer, more complete, or more useful.



