All labs
intermediatePython
Analyse Sales Data with Python
You will learn to
- Import the pandas library and read a CSV file into a DataFrame
- Inspect DataFrame shape, data types, and missing values
- Filter rows based on complex conditions
- Perform data aggregation using groupby()
- Create new derived columns and export results
Before you start
- Basic Python (try Python: Your First Program first)
- The in-browser Jupyter Sandbox (embedded below)
- The sample dataset (download below)
Your workspace
The notebook runs entirely in your browser (JupyterLite). The first load downloads the Python runtime, give it a minute on slow connections, then it is cached. Use the upload arrow in the file panel to add datasets.
Step-by-step
1. Get the dataset
Done when: teki-sales-sample.csv appears in the Jupyter file list2. Import pandas and load data
Done when: A nicely formatted table showing the first five rows appears3. Understand the Data Structure
Done when: You have viewed both the data structure info and the statistical summary4. Filter for high-value transactions
Done when: You have created a new dataframe containing only transactions over $5005. Calculate Revenue by Region
Done when: A Series is printed showing total revenue per region, sorted highest to lowest6. Derive new insights
Done when: A new "price_per_unit" column successfully calculated for every row7. Export your findings
Done when: The enhanced_sales.csv file is generated and visible in the file explorer
Finished every step?
Reflect
- What is the difference between `df.info()` and `df.describe()`?
- Why did we use `ascending=False` when sorting the grouped data?
- What happens if you forget `index=False` when exporting to CSV?