03-cs01-intro

Professor Shannon Ellis

UC San Diego
COGS 137 - Fall 2026

2026-10-05

CS01: Biomarkers of Recent Use

Q&A (10/2)

Q: I’m a little curious how often we’ll be using integer variable types, and what we’ll be using them for. At the moment I don’t see a clear benefit.

A: For our purposes, we won’t. Numerics will work for our needs. The only times one would really want to ensure an integer are: 1) if a function input/code requires an integer or 2) when you want to be very explicit about counts or indices, which will always be integers.

Q: curious how loops would work but i presume it will be largely irrelevant

A: Ah, loops! They have a unique syntax in R. I’ll include one here for reference, but you’re right that we won’t be looping often.

for (i in 1:5) {
  print(i)
}

Q&A (10/5)

Q: I am curious about what would happen with a vector that uses a boolean with a character string.

A: It would coerce to a string! c(FALSE, "shan")

Q: It seems like you can’t get any errors in R, is there a way to get an error.

A: Yes! Lots of errors - we’ll see some in class. But, one note is that they are red….but so are warnings and other output….so just b/c you see red doesn’t mean that there is an error. But, if you’re thinking about dividing by zero giving Inf, you’re right that that would produce an error in other programming languages. ((FALSE, "shan") would return an error…b/c it’s missing the c before the open parentheses if you want ot see one)

Course Announcements

Due Dates:

  • 🔘 Lecture Participation survey “due” after class
  • Lab01 due Friday (we’ll cover material on Wed, so I wouldn’t start this yet)
  • HW01 due Sunday (can get started; we’ll finish content in next set of lecture notes)

Agenda

  • Background
  • Data Intro
  • Paper Results
  • Wrangle

Background

Case Studies in COGS 137

  • Question + Data Provided
  • Wrangling + Analysis Discussed/Presented in Class
  • In groups:
    • reproduce the analysis
    • extend the analysis

Note: we’re discussing this before you know who your CS01 group mates are. This is because everything here is what everyone will need to know. Also, everyone will individually be exploring the data in lab on Thursday…and I want everyone to have the background and work individually to understand the data.

Deliverables

  • Full qmd report (explanations, text and story matter!)
  • General Audience communication

CS01: Biomarkers of Recent Use

  • Focuses on Data Wrangling and EDA
  • Uses real research data from a collaboration w/ Rob Fitzgerald’s group

Motor Vehicle Accidents (MVAs)

  • 2/3 of US trauma center admissions are due to MVAs
  • ~60% of such patients testing positive for drugs or alcohol
  • Alcohol and cannabis are most frequently detected

Source: Hartman and Huestis, 2013

Legalization of Marijuana

  • Federally: recreational use still illegal (Schedule I); state-licensed medical marijuana moved to Schedule III in April 2026
  • Medical use legal in ~40 states + DC
  • Recreational use legal in 24 states + DC (including CA)

Increased roadside surveys

  • 25% increase in use nationwide from 2002 to 2015 (survey)
  • THC detection in drivers increased by 48% from 2007 to 2014
  • Increased prevalence of consumption -> possible intoxication -> possible impaired driving -> public health concern

DUI of Alcohol (DUIA)

  • The science is there. Don’t do it.
  • DUIA has decreased since the 1970s
    • % of nighttime, weekend drivers testing over the legal limit (BAC > 0.08 g/dL) decreased from 7.5% (1973) to 2.2% (2007) (NHTSA roadside survey)

DUI of Cannabis

  • In a 2007 survey, 16.3% of nighttime drivers were drug-positive (NHTSA roadside survey)
    • 8.6% of these tested positive for THC
  • Experimental and cognitive studies suggest cannabis-induced impairment increases risk of motor vehicle crashes:

Evidence suggests recent smoking and/or blood THC concentrations 2–5 ng/mL are associated with substantial driving impairment, particularly in occasional smokers. (Hartman & Huestis, Clinical Chemistry, 2013)

Roadside Detection

  • per se laws: “a driver is deemed to have committed an offense if THC is detected at or above a pre-determined cutoff” (Traffic Injury Prevention, 2021)
  • Defining cutoffs for safe driving is difficult
  • THC concentration differs by:
    • “smoking topography” (time to smoke; number of puffs)
    • frequency of use
    • route of ingestion

Per se laws today

As of October 2025 (GHSA):

  • 18 states have zero tolerance or per se cannabis laws
    • Zero tolerance (any detectable amount): 10 states for THC or its metabolites; 4 states for THC only
    • Numeric per se limits: 4 states, 2–5 ng/mL THC in blood (e.g., Ohio 2 ng/mL; Illinois, Montana, Washington 5 ng/mL)
  • Colorado: “permissible inference”: ≥5 ng/mL THC in blood lets a court infer impairment (driver can rebut)
  • Most other states: the prosecution must show the driver was actually impaired

❓ Why might a zero tolerance law for THC metabolites be especially hard on frequent users?

Metabolism

  • peak blood concentrations occur during smoking, then drop rapidly (Schwope et al., 2012)
  • subjective ‘high’ persists for several hours, varies greatly between individuals
  • THC concentrations remain detectable in frequent users longer than occasional users (Bergamaschi et al., 2013)
  • THC and certain metabolites can be detected in blood for weeks to months after use and do not necessarily indicate impairment

Detection

Various approaches:

  1. Detect impairment (officers detect DUIC)
  2. Detect recent use (test for compounds)
  3. Combine recent use + impairment

Focus here: Can we identify a biomarker of recent use?

  • recent use: defined here as within 3h
  • testing THC and metabolites in blood, oral fluid (OF), and breath

Aside: Case Study Report

  • Your Case study will need a background section
  • It can use/summarize/paraphrase the information here (you should cite the source, not me)
  • But, you’re not limited to this information
  • You are allowed/encouraged to dig deeper, include what’s most important, add to, remove, etc.
  • There are a lot of citations in this section - go ahead and peruse them/others/use references in these papers

Question

Which compound, in which matrix, and at what cutoff is the best biomarker of recent use?

The Data

Participants

  • placebo-controlled, double-blinded, randomized study
  • recruited:
    • volunteers 21-55 y/o
    • had a driver’s license
    • self-reported cannabis use >= 4x in the past month
  • Participants were:
    • compensated
    • medically evaluated (for safety)
    • asked to refrain from use for 2d prior to participation
    • exclusion criteria: OF THC concentration ≥5 ng/mL on day of study (n=7)
  • Study included 191 participants

Demographics

Source: Hoffman et al.

Experimental Design

Participants were:

  • randomly assigned to receive a cigarette containing placebo (0.02%), or 5.9% or 13.4% THC
  • Blood, OF and breath were collected prior to smoking
  • smoked a 700 mg cigarette ad libitum within 10 min, with a minimum of four puffs.
  • After smoking, 4 additional OF and breath and 8 blood collections were completed at time points up to ∼6h from the start of smoking.
  • Participants ate and drank water between collections, although not within 10 min of OF collection.

Timeline

Source: Fitzgerald et al.

Consumption

Source: Hoffman et al.

Topography

Source: Hoffman et al.

What do we recall?

  • Summarize what we know about Detection of Impairment for DUIC.
  • Summarize what experiment was carried out.
  • Summarize what we know about the data so far.

Subjective Highness

Source: Hoffman et al.

Our Datasets

Three matrices:

  • Blood (WB): 8 compounds; 190 participants; 9 collection timepoints (T1–T5B)
  • Oral Fluid (OF): 7 compounds; 192 participants; 5 timepoints (T1–T5A)
  • Breath (BR): 1 compound (THC); 191 participants; 5 timepoints (T1–T5A)

❓ One study…so why 190, 192, and 191 participants?

Variables

  • ID | participant identifier
  • Treatment | Placebo, 5.90%, 13.40%
  • Group | frequent vs. occasional user…but labeled differently across files:
    • WB: “Frequent user” / “Occasional user”
    • OF & BR: “Experienced user” / “Not experienced user”
  • FLUID TYPE (WB) / Fluid (OF, BR) | which matrix the sample came from
  • Timepoint | collection code (T1 = before smoking)
  • time.from.start | minutes from the start of smoking (negative = before)
  • & measurements for individual compounds
    • ng/mL in blood and oral fluid
    • pg/pad in breath (THC (pg/pad)): not directly comparable to the others!

❓ What will we need to fix before we can combine these three files?

The Data

You’ll have access once your groups/repos are created…(today I want people to follow along; there will be time to try on your own soon!)

Important

These data are for our use only and not to be shared widely, so this case study cannot be put in your portfolio, but CS02 and your final project can!

WB <- read_csv("data/Blood.csv")
BR <- read_csv("data/Breath.csv")
OF <- read_csv("data/OF.csv")

First Look at the data

# column names, types, and missing values (no actual values, since these data are private)
data_overview <- function(df) {
  tibble(column    = names(df),
         type      = map_chr(df, \(x) class(x)[1]),
         n_missing = map_int(df, \(x) sum(is.na(x))))
}

In class (and in your own repo), glimpse() each data set to see example values too.

First Look at the data (WB)

# A tibble: 14 × 3
   column          type      n_missing
   <chr>           <chr>         <int>
 1 ID              character         0
 2 Treatment       character         0
 3 Group           character         0
 4 FLUID TYPE      character         0
 5 Timepoint       character         0
 6 CBN             numeric           0
 7 CBD             numeric           0
 8 THC             numeric           0
 9 11-OH-THC       numeric           1
10 THC-COOH        numeric           1
11 THC-COOH-Gluc   numeric           1
12 CBG             numeric           0
13 THC-V           numeric           0
14 time.from.start numeric           2

First Look at the data (OF)

# A tibble: 13 × 3
   column          type      n_missing
   <chr>           <chr>         <int>
 1 ID              character         0
 2 Treatment       character         0
 3 Group           character         0
 4 Fluid           character         0
 5 Timepoint       character         0
 6 CBN             numeric           0
 7 CBD             numeric           0
 8 THC             numeric           0
 9 11-OH-THC       numeric          12
10 CBG             numeric           0
11 THC-V           numeric           0
12 THCA-A          numeric          20
13 time.from.start numeric           0

First Look at the data (BR)

# A tibble: 7 × 3
  column          type      n_missing
  <chr>           <chr>         <int>
1 ID              character         0
2 Treatment       character         0
3 Group           character         0
4 Fluid           character         0
5 Timepoint       character         0
6 THC (pg/pad)    numeric           0
7 time.from.start numeric           0

Analysis

Where We’re Headed…

Results from: Hubbard et al (2021) Biomarkers of Recent Cannabis Use in Blood, Oral Fluid and Breath (Journal of Analytical Toxicology)

Fig 1: Pre-smoking

Fig 2: Sensitivity and Specificity

Fig 3: Cross-compound relationship

Fig 4: Cutoffs

Fig 5: Youden

…and if there’s time PPV and Accuracy post 3h

What Came After

Source: Fitzgerald et al.

Recap

  • Could you summarize/explain background presented?
  • Could you summarize the experiment that was done?
  • Could you describe the datasets? (variables, observations, values, etc.)