Lecture 17 - Cloud Computing II
us-east-1 outageCartoon by Forrest Brazeal
chmod worksapt and copy files with scpaws command linet3.microSource: DataCamp
ls, cd, pwd, and so onsudo runs a command as root, the superuserapt (Advanced Package Tool) installs software
apt update refreshes the package listapt install installs a packagesudo apt install python3 installs Python 3sudo apt install python3-pip installs pipbash terminal that Mac and Linux users have, so every command in this course worksUbuntu in the Start menu/home/<your-linux-name>, also written ~C:, which Linux sees as /mnt/c/Downloads folder/mnt/c/Users/<your-windows-name>/Downloads/chmod. They always look open to everyoneWARNING: UNPROTECTED PRIVATE KEY FILE!ls /mnt/c/Users/Linux in the left panel# 1. Copy the key from Downloads to your Linux home
cp /mnt/c/Users/<your-windows-name>/Downloads/datasci350.pem ~/
# 2. Make it readable only by you
chmod 400 ~/datasci350.pem
# 3. Check: the line should start with -r--------
ls -l ~/datasci350.pem
# 4. Connect as usual
ssh -i ~/datasci350.pem ubuntu@<your-public-dns>The EC2 Dashboard
Name and tags, such as datasci350Application and OS Images, choose UbuntuUbuntu Server 26.04 LTS, marked Free tier eligible64-bit (x86) as the architectureSummary panel on the right lists every choice before you launchThe launch wizard. Click the image to read the Summary panel
t3.micro, labelled Free tier eligiblet3.micro, t3.small, t4g.micro, t4g.small, c7i-flex.large and m7i-flex.larget3.micro has 2 vCPUs and 1 GiB of memory, and costs 0.0104 USD an hour on LinuxThe blue Free tier note in the Summary panel describes the old offer: 750 hours a month of t2.micro for a year. That offer ended for accounts opened after 15 July 2025. Your account runs on credits, and t2.micro is not on your list
Create a new key pair and name it, e.g. datasci350ED25519 and the .pem file formatCreate key pair and save the file somewhere safeNetwork settings, tick Allow HTTPS traffic and Allow HTTP traffic from the internetConfigure storage, you can resize the Root volumegp3 is the default type and works for most usesio2 is the fastest and most expensive, for databases that need sub-millisecond latencysc1 and st1 are cheaper hard-disk types for rarely used data. They cannot be the root diskLaunch instance, open the Instances pageConnect to instance to see the instructionsyes once, and SSH remembers the machineubuntu@ip-172-31-13-146exit logs you out and returns you to your own machinegit, python3, pip, boto3, jq, wget, ssh, vim, nano, tmuxThe shell commands work as usual. What is new is the aws command line, already signed in as you:
| Command | What it does |
|---|---|
aws sts get-caller-identity |
Prints which account and user you are signed in as |
aws ec2 describe-instances --output table |
Lists your instances without leaving the browser |
aws ec2 stop-instances --instance-ids i-... |
Stops an instance from the command line |
aws s3 ls |
Lists your S3 buckets |
aws s3 cp report.pdf s3://my-bucket/ |
Copies a file into a bucket |
sudo dnf install <package> |
Installs software for this session only |
dnf where Ubuntu uses aptdnf goes to /usr/bin, which is wiped when the session endsTwo aws commands and what they return. The account number is blanked out, and describe-instances reports one machine running and one already terminated. More commands here
chmod 400 on files in the Windows filesystem, so SSH refuses the keyActions, then Upload file..pem file. It arrives in /home/cloudshell-user.chmod 400 your-key.pem.ssh -i ... command the console gave you.sudo apt updatesudo apt upgrade
-y answers yes to every prompt: sudo apt update && sudo apt upgrade -ysudo apt install
sudo apt install python3sudo apt install python3-pipscp means “secure copy”. It moves files over SSH-i stands for “identity file”, your keyscp runs on your own machine, not on the instance:XXXXXX with your public DNS, from the Instances page:~ at the end. It means the home directoryls shows the file and python3 hello.py runs itssh command stops workingt3.micro left on for a month costs about $12: $7.60 compute, $3.65 IP address, $0.64 diskjupyterssh -i "<your-key>.pem" ubuntu@<public-DNS> -L 8000:localhost:8888sudo apt update && sudo apt upgrade -ysudo apt install -y python3 python3-pip jupyter-notebookjupyter-notebook. python3-notebook installs only the library, without the jupyter commandwhich python3, which pip3, which jupyterjupyter notebookhttp://localhost:8000 and paste the token from the terminal
http://localhost:8888/?token=...) and change 8888 to 8000print('Hello, DATASCI350!')Closing SSH stops Jupyter. tmux keeps a session alive after you disconnect, and tmux attach brings you back to it. The instance bills while it runs. More on tmux here
scp: upload a file from your computer. Use it for files that are not online, such as your own data or private codewget: download a file from the internet straight to the instance, such as a file on GitHub. It skips your computersudo apt install -y python3-numpy python3-pandas python3-matplotlib python3-seabornscp. On your local machine, create a weather dataset with this code, or download weather_data.py# weather_data.py
import pandas as pd
import numpy as np
import datetime
# Set seed for reproducibility
np.random.seed(42)
# Generate dates for the past 30 days
dates = pd.date_range(end=datetime.datetime.now(), periods=30).tolist()
dates = [d.strftime('%Y-%m-%d') for d in dates]
# Generate temperature data with some randomness
temp_high = np.random.normal(75, 8, 30)
temp_low = temp_high - np.random.uniform(10, 20, 30)
precipitation = np.random.exponential(0.5, 30)
humidity = np.random.normal(65, 10, 30)
# Create a structured dataset
weather_data = pd.DataFrame({
'date': dates,
'temp_high': temp_high,
'temp_low': temp_low,
'precipitation': precipitation,
'humidity': humidity
})
# Save to a text file
with open('weather_data.txt', 'w') as f:
f.write("# Weather data for the past 30 days\n")
f.write(weather_data.to_string(index=False))
print("Weather data saved to weather_data.txt")Run the script on your local machine: python3 weather_data.py
It creates weather_data.txt with 30 days of weather data
Upload it to your instance with scp, from a local terminal:
scp -i <your-key>.pem weather_data.txt ubuntu@<your-instance-ip>:~/Run ls on the instance to check that it arrived
Method 2: wget. Download the analysis script straight to the instance. Run this on your EC2 instance:
wget https://raw.githubusercontent.com/danilofreire/datasci350/main/lectures/lecture-17/weather_analysis.py (one line)wget already downloaded it# weather_analysis.py
import pandas as pd
import matplotlib.pyplot as plt
import seaborn as sns
from io import StringIO
# Read the weather data
with open('weather_data.txt', 'r') as f:
lines = f.readlines()
# Skip the header comment
data_str = ''.join(lines[1:])
df = pd.read_csv(StringIO(data_str), sep=r'\s+')
# Print basic statistics
print("Weather Data Analysis:")
print("=====================")
print(f"Number of days: {len(df)}")
print(f"Average high temperature: {df['temp_high'].mean():.1f}°F")
print(f"Average low temperature: {df['temp_low'].mean():.1f}°F")
print(f"Maximum temperature: {df['temp_high'].max():.1f}°F on {df.loc[df['temp_high'].idxmax(), 'date']}")
print(f"Minimum temperature: {df['temp_low'].min():.1f}°F on {df.loc[df['temp_low'].idxmin(), 'date']}")
print(f"Days with precipitation > 1 inch: {len(df[df['precipitation'] > 1])}")
# Create a visualisation
sns.set_style("whitegrid")
fig, ax1 = plt.subplots(figsize=(12, 6))
# Plot temperature range on the left axis
ax1.fill_between(df['date'], df['temp_low'], df['temp_high'], alpha=0.3, color='skyblue')
ax1.plot(df['date'], df['temp_high'], marker='o', color='red', label='High Temp')
ax1.plot(df['date'], df['temp_low'], marker='o', color='blue', label='Low Temp')
ax1.set_ylabel('Temperature (°F)')
# Add precipitation as bars on a second axis, on the right
ax2 = ax1.twinx()
ax2.bar(df['date'], df['precipitation'], alpha=0.3, color='navy', width=0.5, label='Precipitation')
ax2.set_ylabel('Precipitation (inches)', color='navy')
ax2.tick_params(axis='y', labelcolor='navy')
ax2.grid(False)
# One legend for both axes, and every fifth date on the x-axis
lines1, labels1 = ax1.get_legend_handles_labels()
lines2, labels2 = ax2.get_legend_handles_labels()
ax1.legend(lines1 + lines2, labels1 + labels2, loc='upper left')
ax1.set_xticks(df['date'][::5])
ax1.tick_params(axis='x', labelrotation=45)
ax1.set_title('30-Day Weather Report: Temperature Range and Precipitation', fontsize=16)
fig.tight_layout()
# Save the figure
fig.savefig('weather_analysis.png')
print("Analysis complete. Results saved to 'weather_analysis.png'")weather_data.txt (via scp) and weather_analysis.py (via wget)python3 weather_analysis.py, or run it in Jupyterscp, from a local terminal:
scp -i <your-key>.pem ubuntu@<your-instance-ip>:~/weather_analysis.png ./Recap: scp moves files between your computer and the instance. wget downloads files from the internet straight to the instance
You’ve just completed a full data analysis workflow in the cloud 🎉
Data scientists use this same workflow for larger datasets and more complex analyses!
apt and ran Jupyter through a forwarded portscp and wgetaws command lineGET and POST, and status codesrequests, the library that does all this in three lines of code.env file in Lecture 14 was for an APIBefore then:
The final project is out. Groups of three to four, due 8 December 2026. You pull data from a web API (the World Bank by default, or another one you clear with me) and submit everything as a Docker container. Start from the starter repository: https://github.com/danilofreire/datasci350-project-starter