Welcome to Data Science Computing!

Lecture 01: Introduction

Danilo Freire

Department of Data and Decision Sciences
Emory University

Welcome to DATASCI 350:
Data Science Computing! 🎉

Lecture overview

Today’s agenda

  • Introduction
  • Motivation
  • Class logistics
  • Computer setup

Course materials

Course website: https://danilofreire.github.io/datasci350

Course repository: https://github.com/danilofreire/datasci350

The syllabus, the schedule and all the lecture slides are available on our website

The assignments, project guidelines and study guides are available on our GitHub repository

We will use Canvas for course administration, including submitting assignments, accessing grades, and receiving announcements

Please take some time to get to know all three platforms, and reach out if you have any questions 😉

Note

Please remember to check the course repository regularly for updates and announcements!

Nice to meet you! 😊

Instructor

A bit about me

Visiting Assistant Professor in the Department of Data and Decision Sciences

MA from the Graduate Institute Geneva, PhD from King’s College London, Postdoc at Brown University, Senior Lecturer at the University of Lincoln, UK

Research interests: computational social science, experimental methods, policy evaluation, political violence, organised crime

What about you? (time permitting!)

Now it’s your turn! 😉

Please introduce yourself! 👋

Tell us your name, your major, one thing you really like, and something we don’t know about your city or country! 🌍

My teaching philosophy

  • I love teaching and aim to make learning fun
  • Classes where students participate are the best!
  • Hands-on activities help you learn better
  • I am always available to help and answer questions. And I mean it
  • Your feedback helps me improve my teaching. Please let me know what is working and what is not

Teaching assistants

  • The teaching assistants for this course will be announced soon
  • They will be available to help you with assignments, quizzes, and any questions you may have about the course material
  • We are all here to help you! So feel free to ask questions during class, office hours, or via email 😃

Office hours

What for and what not for

  • What office hours are meant for:
    • Applying tools in practice
    • Discussion of issues related to the assignments
    • Boosting your knowledge of data science
  • What these sessions are not meant for:
    • Solving the assignments for you
    • Taking care of developing your coding skills

Class etiquette

  • Coding can be tough and push you out of your comfort zone. If the course pace is too fast, let me know. I expect your commitment, but I do not want anyone to fail
  • You are all keen on data science, but your backgrounds vary. That is great! Some sessions might be more engaging than others. If you are bored, help others or explore new data science areas
  • Always be respectful to each other
  • Ask questions whenever you need to!

Motivation:
What is data science? 👨🏻‍💻👩🏼‍💻

An old classic

An old classic

An old classic

An old classic

An old classic

Our focus!

Rise of the digital information age

Social media data

New data formats

Survey data

Cheap computing power

And what can we do with all these data? 🤔

Scraping the web for social research

Tackling social problems

Reducing hate speech

Monitoring the effects of climate change on health

Calling bullsh*t when you see it

Learn not to be fooled by

  • big data
  • garbage data
  • garbage models
  • weird samples
  • claims of generality
  • implausibly large effect sizes
  • overfitted models

And much more…

And much more…

  • Abundance of data available for research and for governments to make better decisions
    • Opportunities for novel research questions
    • New methods to answer longstanding research questions
  • New technologies also have social implications and can raise important policy issues
    • Ethical concerns
    • Use of technology by malicious actors
    • Government use of technology to censor or monitor citizens

Course overview and logistics 📖 📚 💻

Course objectives

  • Use the command line interface to manage files and directories
  • Work with version control systems to track changes in code and collaborate with others
  • Create reproducible reports and presentations
  • Run and call language models from your own code
  • Collect data from web APIs and process it at scale
  • Run your own server in the cloud and package your work with containers

Key focus areas

  • This course centres around three key areas of the modern data science workflow: reliability, reproducibility, and robustness
  • Reliability:
    • Ensures consistency in results across multiple runs
    • Minimises errors in data processing and analysis
    • Supports accurate interpretation of findings
  • Reproducibility:
    • Allows others to verify and build upon your work
    • Enhances the credibility of research outcomes
    • Facilitates long-term preservation of scientific knowledge
  • Robustness:
    • Enables analyses to handle unexpected data variations
    • Improves the stability of results under different conditions
    • Supports the scalability of methods to larger datasets

Key tools

Key tools

  • Git and GitHub for version control and collaboration

Key tools

  • Ollama to run language models on your own laptop, and OpenRouter to call hosted ones from Python

Key tools

Key tools

  • requests for collecting data from web APIs

Key tools

Key tools

Key tools

  • uv and Docker for reproducible environments

Logistics

Course information

  • Schedule: Tuesdays and Thursdays, 5:30 to 6:45 pm, Math & Science Center W301
  • Syllabus: On the website and the repository
  • Teaching assistants: To be announced
  • Office hours: By appointment. Email me and we will find a time 😉

Assignments

How you will be graded

  • Problem sets: Ten of them, due on Thursdays at 11:59 pm (50%)
  • In-class quizzes: Five of them (30%)
  • Final project: Due on the last day of class (20%)
  • Late policy: 10% off per day late
  • Collaboration: You can discuss assignments with your classmates, but you must write your own code and submit your own work. AI is allowed
  • Academic integrity: Please refer to the syllabus for the university’s policy on academic integrity

Set up 💻 🛠️

Software

OS extras

Other tools

  • We will install these together during the course. If you want to get ahead, start here:
  • VS Code: Our main code editor. Download it here
  • Python: We use the Anaconda distribution. Download it here and follow our tutorial
  • Quarto: How we write reports and slides. Download it here
  • Ollama: Runs language models on your own laptop. Download it here
  • uv: A faster package manager, used later for reproducible environments. Instructions here
  • Docker: Containerisation tool. Download it here
  • Free accounts: OpenRouter for Lecture 14, AWS for Lectures 16 and 17. Both are free

Next class

  • Computational literacy: binary and hexadecimal numbers, ASCII and Unicode
  • The early days of computing, from Konrad Zuse to the machines we use today
  • How programming languages evolved, and what compiled and interpreted mean
  • Time for questions about installing the terminal. You will need it in two weeks
  • Please create a GitHub education account if you do not have one 😉

and that’s all for today! 🎉

Questions?

Thank you very much for your attention! 🙏🏻

See you soon! 😊