The essentials
Quick reference
One focused task per row. Jump to the related section for complete, working examples.
| Use | Syntax | Examples |
|---|---|---|
| Create a DataFrame | df = pd.DataFrame({'name': ['Ada', 'Lin'], 'score': [91, 84]}) | View examples |
| Read CSV data | df = pd.read_csv('scores.csv', usecols=['name', 'score']) | View examples |
| Preview rows | preview = df.head(3) | View examples |
| Inspect column types | types = df.dtypes | View examples |
| Select one column | scores = df['score'] | View examples |
| Select by labels | result = df.loc[df.index[:2], ['name', 'score']] | View examples |
| Select by position | result = df.iloc[:2, [0, 2]] | View examples |
| Filter with a mask | high_scores = df.loc[df['score'].ge(90)] | View examples |
| Match allowed values | selected = df.loc[df['team'].isin({'red', 'blue'})] | View examples |
| Query rows | selected = df.query('score >= @minimum and active') | View examples |
| Update matching rows | df.loc[df['score'].ge(90), 'level'] = 'advanced' | View examples |
| Create a derived column | result = df.assign(percent=lambda frame: frame['score'] / frame['possible'] * 100) | View examples |
| Rename columns | result = df.rename(columns={'score': 'points'}) | View examples |
| Sort rows | result = df.sort_values(['team', 'score'], ascending=[True, False]) | View examples |
| Set an index | indexed = df.set_index('student_id', verify_integrity=True) | View examples |
| Restore a column index | flat = indexed.reset_index() | View examples |
A DataFrame aligns labeled columns and rows. Select by labels with loc, by integer positions with iloc, and perform assignment in one operation so the intended object is updated under pandas Copy-on-Write semantics.
Step by step
Detailed examples
Create a table and inspect its contract
DataFrame construction requires column values that can align to a common index. When reading external data, constrain columns and parsing options at the boundary, then inspect shape, labels, and dtypes before transforming values. A small head preview is useful, but it does not prove that the whole dataset follows the inferred schema.
from io import StringIO
import pandas as pd
csv = StringIO('name,team,score,ignored\nAda,red,91,x\nLin,blue,84,y\nSam,red,96,z\n')
df = pd.read_csv(csv, usecols=['name', 'team', 'score'])
print(df.head(3).to_string(index=False))
print(df.dtypes.astype(str).to_dict()) name team score
Ada red 91
Lin blue 84
Sam red 96
{'name': 'str', 'team': 'str', 'score': 'int64'}Note: String dtype display depends on the pandas version and configured dtype backend; validate semantic types rather than relying only on printed dtype names.
Select with labels or integer positions
Bracket selection returns a Series for one column. loc interprets both axes as labels and includes both ends of a label slice, while iloc uses Python-style integer positions and excludes the slice stop. Select both axes explicitly when downstream code depends on a stable table shape.
import pandas as pd
df = pd.DataFrame(
{'name': ['Ada', 'Lin', 'Sam'], 'team': ['red', 'blue', 'red'], 'score': [91, 84, 96]},
index=['a1', 'l2', 's3'],
)
scores = df['score']
by_label = df.loc[['a1', 'l2'], ['name', 'score']]
by_position = df.iloc[:2, [0, 2]]
print(scores.to_dict())
print(by_label.equals(by_position.set_axis(by_label.index))) {'a1': 91, 'l2': 84, 's3': 96}
TrueBuild explicit boolean filters
Boolean masks align by index, so derive them from the same DataFrame unless deliberate alignment is part of the operation. Use parentheses around combined comparisons, isin for membership, and query for readable expressions over valid column names. Values injected into query with @ remain data, but query strings should still be developer-controlled rather than accepted as an authorization boundary.
import pandas as pd
df = pd.DataFrame({
'name': ['Ada', 'Lin', 'Sam', 'Mia'],
'team': ['red', 'blue', 'green', 'red'],
'score': [91, 84, 96, 88],
'active': [True, True, False, True],
})
high_scores = df.loc[df['score'].ge(90), ['name', 'score']]
selected_teams = df.loc[df['team'].isin({'red', 'blue'})]
minimum = 85
active_scores = df.query('score >= @minimum and active')
print(high_scores['name'].tolist())
print(selected_teams['name'].tolist())
print(active_scores['name'].tolist()) ['Ada', 'Sam']
['Ada', 'Lin', 'Mia']
['Ada', 'Mia']Assign without chained indexing
Use one loc assignment to target the original DataFrame. Chained assignment cannot update the parent object under pandas 3 Copy-on-Write and is ambiguous in earlier releases. assign is useful in expression pipelines because it returns a new DataFrame and evaluates callable columns against the DataFrame produced so far.
import pandas as pd
df = pd.DataFrame({
'name': ['Ada', 'Lin', 'Sam'],
'score': [91, 42, 96],
'possible': [100, 50, 120],
})
df.loc[df['score'].ge(90), 'level'] = 'advanced'
result = df.assign(percent=lambda frame: frame['score'] / frame['possible'] * 100)
print(result[['name', 'level', 'percent']].to_string(index=False)) name level percent
Ada advanced 91.0
Lin NaN 84.0
Sam advanced 80.0Note: Do not write df[df['score'] >= 90]['level'] = 'advanced'; that chained operation does not reliably update df.
Make labels and ordering intentional
rename changes labels without touching values. sort_values provides deterministic ordering when all relevant tie-breakers are included. set_index makes a domain key available for label-based alignment; verify_integrity checks uniqueness immediately but has a cost, so use it where duplicate identifiers indicate invalid data. reset_index restores index levels as columns.
import pandas as pd
df = pd.DataFrame({
'student_id': [102, 101, 103],
'team': ['red', 'blue', 'red'],
'score': [91, 84, 96],
})
renamed = df.rename(columns={'score': 'points'})
ordered = renamed.sort_values(['team', 'points'], ascending=[True, False])
indexed = ordered.set_index('student_id', verify_integrity=True)
flat = indexed.reset_index()
print(flat.to_dict('records')) [{'student_id': 101, 'team': 'blue', 'points': 84}, {'student_id': 103, 'team': 'red', 'points': 96}, {'student_id': 102, 'team': 'red', 'points': 91}]Local code tester
Filter and update a DataFrame
Change the scores or threshold and run a Copy-on-Write-safe pandas selection locally.
Press Run to load Python locally.
Sources and further reading
References
Authoritative documentation used to verify and expand this cheat sheet.
Help us improve
Found a typo or missing example?
Tell us what would make this cheat sheet clearer, more complete, or more useful.



