WB <- read_csv("data/Blood.csv")
BR <- read_csv("data/Breath.csv")
OF <- read_csv("data/OF.csv")03-cs01-intro
CS01: Biomarkers of Recent Use
Q&A (10/2)
A: For our purposes, we won’t. Numerics will work for our needs. The only times one would really want to ensure an integer are: 1) if a function input/code requires an integer or 2) when you want to be very explicit about counts or indices, which will always be integers.
A: Ah, loops! They have a unique syntax in R. I’ll include one here for reference, but you’re right that we won’t be looping often.
for (i in 1:5) {
print(i)
}Q&A (10/5)
A: It would coerce to a string! c(FALSE, "shan")
A: Yes! Lots of errors - we’ll see some in class. But, one note is that they are red….but so are warnings and other output….so just b/c you see red doesn’t mean that there is an error. But, if you’re thinking about dividing by zero giving Inf, you’re right that that would produce an error in other programming languages. ((FALSE, "shan") would return an error…b/c it’s missing the c before the open parentheses if you want ot see one)
Course Announcements
Due Dates:
- 🔘 Lecture Participation survey “due” after class
- Lab01 due Friday (we’ll cover material on Wed, so I wouldn’t start this yet)
- HW01 due Sunday (can get started; we’ll finish content in next set of lecture notes)
Agenda
- Background
- Data Intro
- Paper Results
- Wrangle
Background
Case Studies in COGS 137
- Question + Data Provided
- Wrangling + Analysis Discussed/Presented in Class
- In groups:
- reproduce the analysis
- extend the analysis
. . .
Note: we’re discussing this before you know who your CS01 group mates are. This is because everything here is what everyone will need to know. Also, everyone will individually be exploring the data in lab on Thursday…and I want everyone to have the background and work individually to understand the data.
Deliverables
- Full qmd report (explanations, text and story matter!)
- General Audience communication
CS01: Biomarkers of Recent Use
- Focuses on Data Wrangling and EDA
- Uses real research data from a collaboration w/ Rob Fitzgerald’s group
Motor Vehicle Accidents (MVAs)
- 2/3 of US trauma center admissions are due to MVAs
- ~60% of such patients testing positive for drugs or alcohol
- Alcohol and cannabis are most frequently detected
Source: Hartman and Huestis, 2013
Legalization of Marijuana
- Federally: recreational use still illegal (Schedule I); state-licensed medical marijuana moved to Schedule III in April 2026
- Medical use legal in ~40 states + DC
- Recreational use legal in 24 states + DC (including CA)
Increased roadside surveys
- 25% increase in use nationwide from 2002 to 2015 (survey)
- THC detection in drivers increased by 48% from 2007 to 2014
- Increased prevalence of consumption -> possible intoxication -> possible impaired driving -> public health concern
DUI of Alcohol (DUIA)
- The science is there. Don’t do it.
- DUIA has decreased since the 1970s
- % of nighttime, weekend drivers testing over the legal limit (BAC > 0.08 g/dL) decreased from 7.5% (1973) to 2.2% (2007) (NHTSA roadside survey)
DUI of Cannabis
- In a 2007 survey, 16.3% of nighttime drivers were drug-positive (NHTSA roadside survey)
- 8.6% of these tested positive for THC
- Experimental and cognitive studies suggest cannabis-induced impairment increases risk of motor vehicle crashes:
. . .
Evidence suggests recent smoking and/or blood THC concentrations 2–5 ng/mL are associated with substantial driving impairment, particularly in occasional smokers. (Hartman & Huestis, Clinical Chemistry, 2013)
Roadside Detection
- per se laws: “a driver is deemed to have committed an offense if THC is detected at or above a pre-determined cutoff” (Traffic Injury Prevention, 2021)
. . .
- Defining cutoffs for safe driving is difficult
- THC concentration differs by:
- “smoking topography” (time to smoke; number of puffs)
- frequency of use
- route of ingestion
. . .
Per se laws today
As of October 2025 (GHSA):
- 18 states have zero tolerance or per se cannabis laws
- Zero tolerance (any detectable amount): 10 states for THC or its metabolites; 4 states for THC only
- Numeric per se limits: 4 states, 2–5 ng/mL THC in blood (e.g., Ohio 2 ng/mL; Illinois, Montana, Washington 5 ng/mL)
- Colorado: “permissible inference”: ≥5 ng/mL THC in blood lets a court infer impairment (driver can rebut)
- Most other states: the prosecution must show the driver was actually impaired
. . .
- Researchers continue to question whether any blood THC cutoff tracks impairment (Bridging THC Knowledge Gaps, 2025)
❓ Why might a zero tolerance law for THC metabolites be especially hard on frequent users?
Metabolism
- peak blood concentrations occur during smoking, then drop rapidly (Schwope et al., 2012)
- subjective ‘high’ persists for several hours, varies greatly between individuals
- THC concentrations remain detectable in frequent users longer than occasional users (Bergamaschi et al., 2013)
- THC and certain metabolites can be detected in blood for weeks to months after use and do not necessarily indicate impairment
Detection
Various approaches:
- Detect impairment (officers detect DUIC)
- Detect recent use (test for compounds)
- Combine recent use + impairment
. . .
Focus here: Can we identify a biomarker of recent use?
- recent use: defined here as within 3h
- testing THC and metabolites in blood, oral fluid (OF), and breath
Aside: Case Study Report
- Your Case study will need a background section
- It can use/summarize/paraphrase the information here (you should cite the source, not me)
- But, you’re not limited to this information
- You are allowed/encouraged to dig deeper, include what’s most important, add to, remove, etc.
- There are a lot of citations in this section - go ahead and peruse them/others/use references in these papers
Question
Which compound, in which matrix, and at what cutoff is the best biomarker of recent use?
The Data
Participants
- placebo-controlled, double-blinded, randomized study
. . .
- recruited:
- volunteers 21-55 y/o
- had a driver’s license
- self-reported cannabis use >= 4x in the past month
. . .
- Participants were:
- compensated
- medically evaluated (for safety)
- asked to refrain from use for 2d prior to participation
- exclusion criteria: OF THC concentration ≥5 ng/mL on day of study (n=7)
. . .
- Study included 191 participants
Demographics
Source: Hoffman et al.
Experimental Design
Participants were:
- randomly assigned to receive a cigarette containing placebo (0.02%), or 5.9% or 13.4% THC
- Blood, OF and breath were collected prior to smoking
- smoked a 700 mg cigarette ad libitum within 10 min, with a minimum of four puffs.
- After smoking, 4 additional OF and breath and 8 blood collections were completed at time points up to ∼6h from the start of smoking.
- Participants ate and drank water between collections, although not within 10 min of OF collection.
Timeline
Source: Fitzgerald et al.
Consumption
Source: Hoffman et al.
Topography
Source: Hoffman et al.
What do we recall?
- Summarize what we know about Detection of Impairment for DUIC.
- Summarize what experiment was carried out.
- Summarize what we know about the data so far.
Subjective Highness
Source: Hoffman et al.
Our Datasets
Three matrices:
- Blood (WB): 8 compounds; 190 participants; 9 collection timepoints (T1–T5B)
- Oral Fluid (OF): 7 compounds; 192 participants; 5 timepoints (T1–T5A)
- Breath (BR): 1 compound (THC); 191 participants; 5 timepoints (T1–T5A)
❓ One study…so why 190, 192, and 191 participants?
Variables
ID| participant identifierTreatment| Placebo, 5.90%, 13.40%Group| frequent vs. occasional user…but labeled differently across files:- WB: “Frequent user” / “Occasional user”
- OF & BR: “Experienced user” / “Not experienced user”
FLUID TYPE(WB) /Fluid(OF, BR) | which matrix the sample came fromTimepoint| collection code (T1 = before smoking)time.from.start| minutes from the start of smoking (negative = before)- & measurements for individual compounds
- ng/mL in blood and oral fluid
- pg/pad in breath (
THC (pg/pad)): not directly comparable to the others!
. . .
❓ What will we need to fix before we can combine these three files?
The Data
You’ll have access once your groups/repos are created…(today I want people to follow along; there will be time to try on your own soon!)
These data are for our use only and not to be shared widely, so this case study cannot be put in your portfolio, but CS02 and your final project can!
First Look at the data
# column names, types, and missing values (no actual values, since these data are private)
data_overview <- function(df) {
tibble(column = names(df),
type = map_chr(df, \(x) class(x)[1]),
n_missing = map_int(df, \(x) sum(is.na(x))))
}In class (and in your own repo), glimpse() each data set to see example values too.
First Look at the data (WB)
# A tibble: 14 × 3
column type n_missing
<chr> <chr> <int>
1 ID character 0
2 Treatment character 0
3 Group character 0
4 FLUID TYPE character 0
5 Timepoint character 0
6 CBN numeric 0
7 CBD numeric 0
8 THC numeric 0
9 11-OH-THC numeric 1
10 THC-COOH numeric 1
11 THC-COOH-Gluc numeric 1
12 CBG numeric 0
13 THC-V numeric 0
14 time.from.start numeric 2
First Look at the data (OF)
# A tibble: 13 × 3
column type n_missing
<chr> <chr> <int>
1 ID character 0
2 Treatment character 0
3 Group character 0
4 Fluid character 0
5 Timepoint character 0
6 CBN numeric 0
7 CBD numeric 0
8 THC numeric 0
9 11-OH-THC numeric 12
10 CBG numeric 0
11 THC-V numeric 0
12 THCA-A numeric 20
13 time.from.start numeric 0
First Look at the data (BR)
# A tibble: 7 × 3
column type n_missing
<chr> <chr> <int>
1 ID character 0
2 Treatment character 0
3 Group character 0
4 Fluid character 0
5 Timepoint character 0
6 THC (pg/pad) numeric 0
7 time.from.start numeric 0
Analysis
Where We’re Headed…
Results from: Hubbard et al (2021) Biomarkers of Recent Cannabis Use in Blood, Oral Fluid and Breath (Journal of Analytical Toxicology)
Fig 1: Pre-smoking
Fig 2: Sensitivity and Specificity
Fig 3: Cross-compound relationship
Fig 4: Cutoffs
Fig 5: Youden
. . .
…and if there’s time PPV and Accuracy post 3h
What Came After
Source: Fitzgerald et al.
Recap
- Could you summarize/explain background presented?
- Could you summarize the experiment that was done?
- Could you describe the datasets? (variables, observations, values, etc.)