DATASCI 350 - Data Science Computing

Lecture 17 - Cloud Computing II

Danilo Freire

Department of Data and Decision Sciences
Emory University

Hello, everyone! 👋

Brief recap 📚

Where we left off

You now have an AWS account

  • Last class: why rent servers instead of buying them
  • Animoto and The Washington Post showed what renting makes possible
  • IaaS, PaaS, SaaS, FaaS and the main AWS services
  • The October 2025 us-east-1 outage
  • Your free plan account and zero-spend budget
  • Amazon Textract on a scanned document
  • So far, everything happened in the browser. Today we use the terminal (as always!) 😂

Cartoon by Forrest Brazeal

Today’s plan

  • Launch an EC2 instance from the AWS console
  • Connect to it with SSH and a key pair
  • Windows users: keep the key in WSL, where chmod works
  • Install software with apt and copy files with scp
  • Stop and terminate instances, so the bill stays at zero
  • CloudShell: a terminal in the browser, with the aws command line
  • Activity 01: run Jupyter on the instance
  • Activity 02: analyse data in the cloud (time permitting)

EC2 from your local machine 🖥️

EC2 instances in a nutshell

  • EC2 is a virtual server in the cloud, usually running Ubuntu or Amazon Linux
  • Virtualisation runs many instances on one physical machine
  • Instance types differ in CPU, memory, storage, and network
    • T3 and M7i are general purpose
    • C7i is compute-optimised, and R7i is memory-optimised
  • Six types are free tier eligible. We use t3.micro

Source: DataCamp

Ubuntu Linux

  • Linux is an open-source operating system
  • It runs most of the servers on the internet
  • It is free to install, and you can change any part of it
  • Ubuntu is the most widely used Linux distribution
  • Google, Amazon, and Microsoft all offer Ubuntu instances
  • WSL runs Linux inside Windows, Ubuntu by default
  • Ubuntu uses bash as its default login shell
  • So the commands from the shell module work on Ubuntu

Ubuntu server

  • Ubuntu has a GUI, but we will use only the terminal
  • A bare-bones instance is cheaper and faster than a full desktop
  • The terminal has everything you need to manage it
  • The commands are the same: ls, cd, pwd, and so on
  • sudo runs a command as root, the superuser
  • Root has full control over the system, so be careful with it!
  • apt (Advanced Package Tool) installs software
    • apt update refreshes the package list
    • apt install installs a package
  • sudo apt install python3 installs Python 3
  • sudo apt install python3-pip installs pip

Windows Subsystem for Linux (WSL)

A real Linux system inside Windows

  • WSL runs Linux inside Windows. It installs Ubuntu by default
  • You get the same bash terminal that Mac and Linux users have, so every command in this course works
  • Windows runs it in a small virtual machine and starts it for you when you open the Ubuntu app
  • Tutorial 06 has the installation steps. I hope many of you have done it already! 😊
  • To open it, type Ubuntu in the Start menu
  • WSL has two separate filesystems:
    • Your Linux home, /home/<your-linux-name>, also written ~
    • Your Windows drive C:, which Linux sees as /mnt/c/

WSL and your SSH key

Move the key into your Linux home first

  • Your browser saves the key on the Windows side, in your Downloads folder
  • From WSL, that folder is /mnt/c/Users/<your-windows-name>/Downloads/
  • Files on the Windows drive ignore chmod. They always look open to everyone
  • SSH then refuses the key with WARNING: UNPROTECTED PRIVATE KEY FILE!
  • The fix is to copy the key into your Linux home and set its permissions there
  • Not sure of your Windows user name? Run ls /mnt/c/Users/
  • Your Linux files also appear in Windows Explorer, under Linux in the left panel
# 1. Copy the key from Downloads to your Linux home
cp /mnt/c/Users/<your-windows-name>/Downloads/datasci350.pem ~/

# 2. Make it readable only by you
chmod 400 ~/datasci350.pem

# 3. Check: the line should start with -r--------
ls -l ~/datasci350.pem

# 4. Connect as usual
ssh -i ~/datasci350.pem ubuntu@<your-public-dns>

Step 01: Launching an EC2 instance

  • Now we can launch an EC2 instance!
  • Sign in to the AWS Console and search for EC2
  • On the EC2 Dashboard, click the orange Launch instance button
  • Choose an AMI (Amazon Machine Image). We pick Ubuntu, which is free tier eligible
  • AWS has thousands of AMIs: Windows, macOS, and other Linux distributions
  • Many come ready to run specific software, such as Jupyter, TensorFlow, or Docker

Step 02: Choosing an instance type

  • Type a name under Name and tags, such as datasci350
  • Under Application and OS Images, choose Ubuntu
  • The current release is Ubuntu Server 26.04 LTS, marked Free tier eligible
  • Keep 64-bit (x86) as the architecture
  • Launch 1 instance for now
  • The Summary panel on the right lists every choice before you launch
  • The next slide covers the instance type
  • So far, so good? 😊

The launch wizard. Click the image to read the Summary panel

Which instance type is free?

The console’s own note is out of date

  • Choose t3.micro, labelled Free tier eligible
  • Accounts opened after 15 July 2025 can use six types: t3.micro, t3.small, t4g.micro, t4g.small, c7i-flex.large and m7i-flex.large
  • The dropdown shows the price of every type
  • A t3.micro has 2 vCPUs and 1 GiB of memory, and costs 0.0104 USD an hour on Linux
  • You can change the instance type later

The blue Free tier note in the Summary panel describes the old offer: 750 hours a month of t2.micro for a year. That offer ended for accounts opened after 15 July 2025. Your account runs on credits, and t2.micro is not on your list

The instance type dropdown, with the hourly price under every entry

Step 03: Choose an SSH key pair

  • Next, choose a key pair
  • A key pair is two cryptographic keys that authenticate you to an instance
  • You create it once and use it for all your instances
  • Click Create a new key pair and name it, e.g. datasci350
  • Choose ED25519 and the .pem file format
  • Click Create key pair and save the file somewhere safe
  • You need it to log in over SSH, and AWS cannot send you another copy

Step 04: Check HTTP/HTTPS options

  • Under Network settings, tick Allow HTTPS traffic and Allow HTTP traffic from the internet
  • This lets you reach web servers on your instance
  • SSH (port 22) is open by default
  • We open these ports to anyone for now
  • In production, allow only what you need
  • Why? Security! Every open port is a way in for attackers
  • You can edit the security group later

Step 05: Configure storage

  • Under Configure storage, you can resize the Root volume
  • The default 8 GB is enough for today. A larger disk spends more credits
  • You can add more volumes if you need them
  • gp3 is the default type and works for most uses
  • io2 is the fastest and most expensive, for databases that need sub-millisecond latency
  • sc1 and st1 are cheaper hard-disk types for rarely used data. They cannot be the root disk

  • Now you just have to click on Launch instance! 🚀

Step 06: Log in to the EC2 instance through SSH

  • After you click Launch instance, open the Instances page
  • Click Connect to instance to see the instructions

  • Choose the SSH client tab
  • Open a terminal in the folder where you saved the key
  • The console writes both commands for you under How to connect:
    • Step 3: chmod 400 "your-key.pem"
    • Step 4: ssh -i "your-key.pem" ubuntu@your-public-dns

And you are in! 🎉

Welcome to the cloud! 🌥️

  • The first connection asks you to confirm the host fingerprint
  • Type yes once, and SSH remembers the machine
  • The banner shows what you rented: Ubuntu 26.04 LTS, a 6.61 GB disk that is 30% full, 118 processes, and no other users
  • The prompt changes to ubuntu@ip-172-31-13-146
  • Every command now runs on a computer in Virginia
  • exit logs you out and returns you to your own machine

Check the instance details on the AWS Console

  • Click the instance ID, or Instances on the left
  • You can see its public IP address, instance type, security group, and key pair
  • You can also stop, terminate, or reboot it here (more on this later)

AWS CloudShell

A terminal in the browser, already signed in

  • Open it from the bottom left of the console. Nothing to install
  • Amazon Linux 2023: 1 vCPU, 2 GiB of RAM, 1 GB of saved storage
  • Free: you pay only for the resources you use from it
  • The AWS CLI is already signed in as you
  • Also there: git, python3, pip, boto3, jq, wget, ssh, vim, nano, tmux
  • It runs bash, so your usual commands work
  • Sessions end after 20 to 30 idle minutes, or 12 hours
  • Unused home directories are deleted after 120 days

The panel you meet on first launch, stating the tools and the 1 GB of storage per region

  • I recommend your local terminal and editor. Use CloudShell as a fallback

CloudShell commands

What is actually new

The shell commands work as usual. What is new is the aws command line, already signed in as you:

Command What it does
aws sts get-caller-identity Prints which account and user you are signed in as
aws ec2 describe-instances --output table Lists your instances without leaving the browser
aws ec2 stop-instances --instance-ids i-... Stops an instance from the command line
aws s3 ls Lists your S3 buckets
aws s3 cp report.pdf s3://my-bucket/ Copies a file into a bucket
sudo dnf install <package> Installs software for this session only
  • Amazon Linux uses dnf where Ubuntu uses apt
  • Software installed with dnf goes to /usr/bin, which is wiped when the session ends
  • Keep anything you need in your home directory

Two aws commands and what they return. The account number is blanked out, and describe-instances reports one machine running and one already terminated. More commands here

CloudShell as a way back in

When your own terminal will not cooperate

  • The WSL problem: Windows ignores chmod 400 on files in the Windows filesystem, so SSH refuses the key
  • CloudShell avoids this, because the key goes to a Linux machine:
  1. Open CloudShell from the console.
  2. Choose Actions, then Upload file.
  3. Select your .pem file. It arrives in /home/cloudshell-user.
  4. Run chmod 400 your-key.pem.
  5. Run the same ssh -i ... command the console gave you.
  • The instance does not care that the connection came from a browser tab

The whole thing from a browser tab: the upload confirmation, chmod 400, and an Ubuntu 26.04 prompt on the instance

  • CloudShell times out after 20 to 30 idle minutes, and your SSH connection goes with it

Step 07: Update and install software

  • First, refresh the package list: sudo apt update
  • Then install the latest updates: sudo apt upgrade
    • -y answers yes to every prompt: sudo apt update && sudo apt upgrade -y
  • Install software with sudo apt install
    • Python 3: sudo apt install python3
    • pip: sudo apt install python3-pip
  • You can install Jupyter and anything else the same way

sudo apt update && sudo apt upgrade -y && sudo apt install -y python3 python3-pip

pip install gives an externally-managed-environment error on Ubuntu. Use sudo apt install python3-<package> or a virtual environment

Adding files to your instance

  • scp means “secure copy”. It moves files over SSH
  • -i stands for “identity file”, your key
  • The first path is the source, and the second is the destination
# laptop to instance
scp -i key.pem myfile ubuntu@public-dns:~

# instance to laptop
scp -i key.pem ubuntu@public-dns:~/myfile .
  • The . in the second command means “the folder I am in now”
  • scp can also copy between two remote servers
  • Windows users: WinSCP and PuTTY do the same with a drag-and-drop window
  • Open a new local terminal. scp runs on your own machine, not on the instance:
echo 'print("Hello, DATASCI350!")' > hello.py
scp -i datasci350.pem hello.py ubuntu@XXXXXX.compute-1.amazonaws.com:~

  • Replace XXXXXX with your public DNS, from the Instances page
  • Keep the :~ at the end. It means the home directory
  • In your SSH session, ls shows the file and python3 hello.py runs it

Stopping and terminating your instance

  • Stop or terminate instances you are not using
  • Stop pauses it. Compute is not charged, but the disk still is
  • Terminate deletes it, with every file on it
  • A restarted instance gets a new IP address, so your old ssh command stops working
  • A t3.micro left on for a month costs about $12: $7.60 compute, $3.65 IP address, $0.64 disk
  • Stop, terminate, or reboot from the Instances page
  • Let’s terminate ours now! 🛑

The Instance state menu. Terminate (delete) instance is the last entry

Afterwards both instances read Terminated, and the Public IPv4 DNS column is empty. The address is gone

Now it is your turn! 🚀

Activity 01

  1. Launch an EC2 instance called jupyter
  2. Connect with port forwarding:
    • ssh -i "<your-key>.pem" ubuntu@<public-DNS> -L 8000:localhost:8888
    • This forwards port 8888 on the instance to port 8000 on your laptop
  3. Install Python, pip, and Jupyter:
    • sudo apt update && sudo apt upgrade -y
    • sudo apt install -y python3 python3-pip jupyter-notebook
    • Note the hyphen in jupyter-notebook. python3-notebook installs only the library, without the jupyter command
  4. Check: which python3, which pip3, which jupyter
  1. Start Jupyter: jupyter notebook
  2. Open http://localhost:8000 and paste the token from the terminal
    • Or click the terminal link (http://localhost:8888/?token=...) and change 8888 to 8000
  3. Create a notebook and run print('Hello, DATASCI350!')
  4. Do not terminate the instance yet. We use it again later

Closing SSH stops Jupyter. tmux keeps a session alive after you disconnect, and tmux attach brings you back to it. The instance bills while it runs. More on tmux here

Activity 01

Activity 01

Another one? 🚀

Activity 02

  • Now let’s analyse some data on the instance
  • We practise two ways to get files onto it:
    • Method 1, scp: upload a file from your computer. Use it for files that are not online, such as your own data or private code
    • Method 2, wget: download a file from the internet straight to the instance, such as a file on GitHub. It skips your computer
  • First, install the packages on your instance:
    • sudo apt install -y python3-numpy python3-pandas python3-matplotlib python3-seaborn
  • Method 1: scp. On your local machine, create a weather dataset with this code, or download weather_data.py
# weather_data.py
import pandas as pd
import numpy as np
import datetime

# Set seed for reproducibility
np.random.seed(42)

# Generate dates for the past 30 days
dates = pd.date_range(end=datetime.datetime.now(), periods=30).tolist()
dates = [d.strftime('%Y-%m-%d') for d in dates]

# Generate temperature data with some randomness
temp_high = np.random.normal(75, 8, 30)
temp_low = temp_high - np.random.uniform(10, 20, 30)
precipitation = np.random.exponential(0.5, 30)
humidity = np.random.normal(65, 10, 30)

# Create a structured dataset
weather_data = pd.DataFrame({
    'date': dates,
    'temp_high': temp_high,
    'temp_low': temp_low,
    'precipitation': precipitation,
    'humidity': humidity
})

# Save to a text file
with open('weather_data.txt', 'w') as f:
    f.write("# Weather data for the past 30 days\n")
    f.write(weather_data.to_string(index=False))
    
print("Weather data saved to weather_data.txt")

Activity 02

  • Run the script on your local machine: python3 weather_data.py

  • It creates weather_data.txt with 30 days of weather data

  • Upload it to your instance with scp, from a local terminal:

    • scp -i <your-key>.pem weather_data.txt ubuntu@<your-instance-ip>:~/
  • Run ls on the instance to check that it arrived

  • Method 2: wget. Download the analysis script straight to the instance. Run this on your EC2 instance:

  • wget https://raw.githubusercontent.com/danilofreire/datasci350/main/lectures/lecture-17/weather_analysis.py (one line)
  • The code is below for reference. You do not need to copy it, because wget already downloaded it
# weather_analysis.py
import pandas as pd
import matplotlib.pyplot as plt
import seaborn as sns
from io import StringIO

# Read the weather data
with open('weather_data.txt', 'r') as f:
    lines = f.readlines()

# Skip the header comment
data_str = ''.join(lines[1:])
df = pd.read_csv(StringIO(data_str), sep=r'\s+')

# Print basic statistics
print("Weather Data Analysis:")
print("=====================")
print(f"Number of days: {len(df)}")
print(f"Average high temperature: {df['temp_high'].mean():.1f}°F")
print(f"Average low temperature: {df['temp_low'].mean():.1f}°F")
print(f"Maximum temperature: {df['temp_high'].max():.1f}°F on {df.loc[df['temp_high'].idxmax(), 'date']}")
print(f"Minimum temperature: {df['temp_low'].min():.1f}°F on {df.loc[df['temp_low'].idxmin(), 'date']}")
print(f"Days with precipitation > 1 inch: {len(df[df['precipitation'] > 1])}")

# Create a visualisation
sns.set_style("whitegrid")
fig, ax1 = plt.subplots(figsize=(12, 6))

# Plot temperature range on the left axis
ax1.fill_between(df['date'], df['temp_low'], df['temp_high'], alpha=0.3, color='skyblue')
ax1.plot(df['date'], df['temp_high'], marker='o', color='red', label='High Temp')
ax1.plot(df['date'], df['temp_low'], marker='o', color='blue', label='Low Temp')
ax1.set_ylabel('Temperature (°F)')

# Add precipitation as bars on a second axis, on the right
ax2 = ax1.twinx()
ax2.bar(df['date'], df['precipitation'], alpha=0.3, color='navy', width=0.5, label='Precipitation')
ax2.set_ylabel('Precipitation (inches)', color='navy')
ax2.tick_params(axis='y', labelcolor='navy')
ax2.grid(False)

# One legend for both axes, and every fifth date on the x-axis
lines1, labels1 = ax1.get_legend_handles_labels()
lines2, labels2 = ax2.get_legend_handles_labels()
ax1.legend(lines1 + lines2, labels1 + labels2, loc='upper left')
ax1.set_xticks(df['date'][::5])
ax1.tick_params(axis='x', labelrotation=45)
ax1.set_title('30-Day Weather Report: Temperature Range and Precipitation', fontsize=16)
fig.tight_layout()

# Save the figure
fig.savefig('weather_analysis.png')
print("Analysis complete. Results saved to 'weather_analysis.png'")

Activity 02

  • Both files are now on your instance: weather_data.txt (via scp) and weather_analysis.py (via wget)
  • Run the analysis: python3 weather_analysis.py, or run it in Jupyter
  • Download the image to your local machine with scp, from a local terminal:
    • scp -i <your-key>.pem ubuntu@<your-instance-ip>:~/weather_analysis.png ./
  • Open the image on your laptop
  • Don’t forget to terminate your instance when done!

Recap: scp moves files between your computer and the instance. wget downloads files from the internet straight to the instance

Activity 02 Result

When you run the activity, you should get a graph like this one:

Activity 02 Result

You’ve just completed a full data analysis workflow in the cloud 🎉

  1. Created data locally
  2. Uploaded it to the cloud
  3. Processed it on a cloud server
  4. Generated visualisations
  5. Downloaded results to your local machine

Data scientists use this same workflow for larger datasets and more complex analyses!

Conclusion

Summary

What we learned today

  • We launched an EC2 instance and connected over SSH with a key pair
  • We installed software with apt and ran Jupyter through a forwarded port
  • We moved files with scp and wget
  • We met CloudShell and the aws command line
  • We learnt to stop and terminate instances, to keep the bill at zero
  • Lambda runs a function with no server to manage. It is the easiest service to try next
  • S3, RDS, and SageMaker cover files, databases, and machine learning
  • The same workflow works for data too big for a laptop: upload it, rent a machine, run the code, and download the results

Next class

Getting data from the web

  • We leave the cloud and look at how data gets from the internet into your analysis
  • What an API is, with a client and a server on either side
  • HTTP: URLs, query strings, GET and POST, and status codes
  • JSON, the format most APIs answer in, and how it maps to a Python dictionary
  • requests, the library that does all this in three lines of code
  • You have done this before: the API key in your .env file in Lecture 14 was for an API

Before then:

  1. Terminate every instance you started today
  2. Open Billing and Cost Management and check that the total is zero
  3. Read the final project instructions here: https://github.com/danilofreire/datasci350/blob/main/project/project-instructions.pdf

The final project is out. Groups of three to four, due 8 December 2026. You pull data from a web API (the World Bank by default, or another one you clear with me) and submit everything as a Docker container. Start from the starter repository: https://github.com/danilofreire/datasci350-project-starter

And that’s all for today! 🎉