WUSS 2024 Paper Abstracts

WUSS 2024 will feature more than 100 presentations in several different formats that cover a variety of topics and experience levels.

Note: This information is subject to change. Last updated September 3, 2024.

 

Hands-on Workshops (HOWs)

No.Author(s)Paper Title (click for abstract)

101 Troy Martin Hughes Hands-on Python PDFs: Using the pypdf Library To Programmatically Design, Complete, Read, and Extract Data from PDF Forms Having Digital Signatures
170 Matthew T Slaughter, Isaiah Lankham A hands-on introduction to end-to-end data projects in Python for SAS Programmers
192 Louise S Hadden Utilizing the Medicare Part D Prescriber Public Use Data File as a “Training” Data Set
198 Isaiah Lankham, Matthew T Slaughter Commit early, commit often! A gentle, hands-on introduction to the joy of Git and GitHub
200 Josh Horstman Map It Out: Using SG Attribute Maps for Precise Control of PROC SGPLOT Output
207 Matt Becker Mastering Clinical Trial Reporting with SAS
209 Zeke Torres GIT Intro for SAS Coders that need to GIT started

 

Live Demonstrations

No.Author(s)Paper Title (click for abstract)

121 Mark Jordan Jedi SAS Tricks for DS2 Programmers
126 Mark Jordan Five Easy Ways to Make Your SAS Code Run Faster
128 Mark Jordan Beyond Macro – Data-Driven Programming with CAS in SAS Viya
135 Melodie Rush Handling Missing Values in SAS 9 and SAS Viya
156 Jesse Albert Canchola Correlation of a Composite Self-Reported Symptoms Severity Score with Patient Number of Symptoms as Predictors of SARS-CoV-2 Viral Status/Load using the SAS(r) System
163 Inka Leprince CTRL + S(HARE): A Programmer’s Advice for Branding on Social Media
164 Bart Jablonski Here Comes the Rain (Cloud Plot) Again
165 Chary Akmyradov Streamlining Your Workflow: Creating Portable and Automated SAS Project Folders Using Your SAS Enterprise Guide Project Name
168 Danny R Modlin PROC BGLIMM: The Smooth Transition to Bayesian Analysis
173 Charu Shankar SAS Visual Analytics a low code no code approach to dashboard reporting
178 Sarvar Khamidov, James Joseph The Dataset-JSON Submission Process
184 Christopher Johnson Integration of an R Language Function Within SAS to Perform Conditional Multivariate Saddlepoint Approximation
201 Lisa A Mendez How RANK are your deciles? Using PROC RANK and PROC MEANS to create deciles based on observations and numeric values
206 Matt Becker SAS Viya in Action: Enhancing Life Sciences Data Analysis
211 Casey Smith SAS Studio, SAS Enterprise Guide, SAS Extension for Visual Studio Code: Which should I use?
212 Chris Hemedinger Using VS Code with SAS
213 Chris Hemedinger Working with SAS and Microsoft 365
216 Troy Martin Hughes Who’s Still Bringing That Big Data Energy? Leading a SAS® Conference by Leveraging a 48-Year Longitudinal Analysis of 30,498 Presentations in the SAS User Community
217 Troy Martin Hughes From Word Clouds to Phrase Clouds to Amaze Clouds: A Data-Driven Python Programming Solution To Building Configurable Taxonomies That Standardize, Categorize, and Visualize Phrase Frequency
218 Joshua J Cook, Swann Arp Adams An introduction to Quarto: A Versatile Open-Source Tool for Data Reporting and Visualization
219 Matt Becker SAS-tastic Visuals: Turning Data into Eye Candy!

 

Lively Lectures

No.Author(s)Paper Title (click for abstract)

102 Richann Jean Watson, Louise S Hadden Just Stringing Along: FIND Your Way to Great User-Defined Functions
103 Jane Eslinger Incorporating Macro with PROC REPORT Code
104 Kirk Paul Lafler SAS® Macro Programming Tips and Techniques
108 Kirk Paul Lafler Creating Custom Excel Spreadsheets with Built-in Autofilters Using SAS® Output Delivery System (ODS)
112 David Horvath Introduction to developing Neural Nets
113 David Horvath Leveraging “UNIX Tools” (GNU) for Data Analysis
114 Josh Horstman, Richann Jean Watson From Muggles to Macros: Transfiguring Your SAS® Programs With Dynamic, Data-Driven Wizardry
115 Ryan Paul Lafler, Anna T. K. Wade Developing Artificial and Convolutional Neural Networks with Python’s Keras API for TensorFlow
116 Ryan Paul Lafler Charting Your Organization’s Machine Learning Roadmap
117 Lisa A Mendez, Richann Jean Watson LAST CALL to Get Tipsy with SAS®: Tips for Using CALL Subroutines
118 Scott Burroughs Can You Teach An Old Dog New Tricks?
119 Jesse Albert Canchola Limit of Detection Calculation Methods Comparison for PCR-based Quantitative Studies using the SAS® System
120 Jayanth Iyengar SAS Job Searching and Interviewing tips – Strategies in the Post-Pandemic era
122 Evelyn Larson, Anabeli Chandra Airbnb Consumer Behavior Analysis
123 Melodie Rush How does SAS Viya and Open Source Integrate
124 Bart Jablonski Macro Variable Arrays Made Easy with macroArray SAS package
125 Bart Jablonski, Quentin McMullen Fifty Shades of SAS Programming, or 53 (+3) Syntax Snippets for a Table Look-up Task, or How to Learn SAS by Solving Only One Exercise!
130 Rick Wicklin Matrices and More: An Overview of the SAS/IML Language
131 Rick Wicklin Ten tips for simulating data with SAS
132 Junying Zhang Automation of Submission of Programs in Analysis Data Reviewer’s Guide for FDA and PMDA Submissions
133 Ron Cody A Survey of Some of the Most Useful SAS Functions
136 Iuliana Constantin Navigating Career Crossroads: From Statistical Manager to Entrepreneur
143 Derek Morgan The Essentials of SAS® Dates and Times
144 Derek Morgan PROC SORT (then and) NOW
145 Derek Morgan Time Since Last Dose: Anatomy of a SQL Query
147 Jayanth Iyengar The Everytown Research database: Using SAS® analytic procedures to analyze mass shootings
148 Paul Michael Dorfman Uniform Hashing of Arbitrary Input into Key-Exclusive Segments
149 Joe Madden Putting Power into the Hands of the Programmer with SAS Viya Workbench
150 LeRoy Bessler Bessler’s Principles of Communication-Effective Data Visualization
151 LeRoy Bessler Exploration and Revelation for COVID-19 Data: An Atlas and Other Visual Data Insights
152 Don Henderson The SAS Supervisor
153 Paul Michael Dorfman Array Hashing: Simple, Fast, and Efficient
155 Menaga Guruswamy Ponnupandy Leading with Impact: Building Stronger Programming Teams Through Effective Leadership and Collaboration
157 Junying Zhang Will statistical programmers be replaced by Artificial Intelligence?
159 Mehrnaz Siavoshi Estimating causal effects of health interventions using instrumental variables
160 Sudhir Kedare, Steve Wade iCSR: A Wormhole to Interactive Data Exploration Universe
161 Steve Wade, Sudhir Kedare inspectoR: QC in R? No Problem!
162 Steve Wade, Sudhir Kedare Interactive Data Analysis and Exploration with composR: See the Forest AND the Trees
166 Marckenley Mercie, Ce (Colin) Zhou Statistical Programmers Role for SPM Analysis
169 Danny R Modlin Customizing SAS Studio in Viya: The Next Step
171 Kathryn Doh, Maggie G Tufts Identifying a Skilled Nursing Facility Associated COVID-19 Case in Surveillance Data
172 Danny R Modlin How to Modify SAS 9 Programs to Run in SAS Viya
174 Riley Rutan, Eric C Braga, Xinyu Du X-Ray Image Classification with Neural Networks
175 Eric C Braga, Riley Rutan, Miguel Angel Bravo Martinez del Valle, Xinyu Du, Fernanda Carrillo Flight Delay Analysis and Prediction
179 Sy J Truong Training Domain Specific AI for Clinical Data: Developing an AI Agent to Transform EDC data to CDISC SDTM and then to ADaM Standards
182 Alec Troy Christiansen, Benjamin T Houser Analyzing Goalkeeper Impact in Premier League Football
183 Rachelle Juan A SAS® Macro to Calculate Therapeutic Intensity Score
185 Michael Shealy SAS Viya vs SAS 9: Architecture and Performance For SAS Users / Teaming With Your SAS Architect and Admin For A Fast SAS Viya
186 James Joseph A Framework for Clinical Trial Data Synthesis
187 Michael Williams Simulation in SAS for Estimating Power
188 Chary Akmyradov, Lida Gharibvand Simulating Optimal Sample Sizes for Joyful Canine Jaws Using SAS
189 Stephen Sloan The Business Value of Diversity Combined with Data Science
190 Zheyuan Yu Regression Analysis Made Easy Using SAS® Studio
191 Bocheng Jing, L. Grisell Diaz-Ramirez, John Boscardin Developing a SAS Macro for ISNI: Assessing Missing Data Sensitivity Beyond MAR Assumptions
193 Joshua J Cook, Louise S Hadden, Swann Arp Adams Comparative Examples of Geocoding and Mapping Techniques in SAS® and R Using Centers for Disease Control Survey Data
194 Leon Rod Davoody, Lida Gharibvand Qualitative and Quantitative Data Analysis and Visualization Using Python
196 Yi-Che Chen, Debanik Chakraborty From Keys to Credit: A Deep Dive into Homeownership’s Influence on Loan Quality
199 Josh Horstman Advanced DATA Step Look-Back and Look-Ahead Techniques
202 Steve Black Count, Lag, Retain! Let’s go through it one more time!
203 Jane Eslinger It’s All about the Base—Procedures
205 Kim Wilson Taking the Mystery Out of and Debugging PROC HTTP
208 Zeke Torres Customized LOG/LST name with the SAS® Config File
210 Bill Coar Multiple Imputation in Interim Analyses
214 Chris Hemedinger The Dummy’s Guide to Getting Started with SAS
215 Jim Blum, Jonathan W Duggins Methods and Tools for Publishing SAS Papers, Books, and Other Documentation
220 Bill Coar Introduction to Conditional Power by Example

 

Panel Discussions

No.Author(s)Paper Title (click for abstract)

106 Kirk Paul Lafler, Ryan Paul Lafler, Joshua J Cook, Stephen Sloan, Anna T. K. Wade Benefits, Challenges, and Opportunities with SAS® and Open-Source Software (OSS) Integration
197 Isaiah Lankham, Matthew T Slaughter What comes after the semicolon? How I learned to stop worrying and love the meeting

 

SAS-quatschen Facilitated Discussions

No.Author(s)Paper Title (click for abstract)

154 Lisa A Mendez, Charu Shankar Get a GPS to Navigate Your Skills to Find Career Purpose
195 Lida Gharibvand The Current State of Teaching Biostatistics in Academia
204 Jane Eslinger Deciding What to Write About in the SAS World

  


Abstracts

Hands-on Workshops (HOWs)

101 : Hands-on Python PDFs: Using the pypdf Library To Programmatically Design, Complete, Read, and Extract Data from PDF Forms Having Digital Signatures
Troy Martin Hughes
Thursday, September 05, 2024, 03:00 PM – 04:00 PM, Location: Golden State
Download White Paper (PDF)

The pypdf Python library (https://pypdf.readthedocs.io/en/stable/index.html) facilitates the programmatic creation, completion, cropping, and merging of PDF forms. Form data—including both dynamic text and field values—can be programmatically written to a PDF using pypdf, and data manually entered into PDF form fields by end users can be programmatically extracted and evaluated. With this combination of functionality, pypdf is a powerful tool that can build dynamically generated PDF forms that simplify user completion of forms, as well as subsequent form validation. This text introduces users to the pypdf library, and demonstrates a single use case in which copyright grant forms (CGFs, aka copyright-release or permission-to-publish forms) were automatically generated for the Western Users of SAS Software (WUSS) 2024 conference proceedings. This automation eliminated confusing language and components of the CGF—for example, by removing language specific only to US government employees (unless the author was a US government employee). Thus, in 2023, an author submitting a paper to WUSS had to navigate more than 50 unutilized fields in the CGF!! Moreover, an author completing the WUSS 2023 CGF had to enter his name, paper number, paper title, job title, and organization—information that the conference already had collected and which it should have been using to prepopulate forms for authors. Thus, the revised form and process now only requires each author to digitally sign the CGF; errors are eliminated and efficiency is maximized. This solution was developed and run on Python 3.11.

170 : A hands-on introduction to end-to-end data projects in Python for SAS Programmers
Matthew T Slaughter, Isaiah Lankham
Thursday, September 05, 2024, 04:00 PM – 05:30 PM, Location: Golden State
Download Slide Presentation (PDF)

Are you interested in learning the world’s most popular programming language, but aren’t sure where to get started?

In this hands-on workshop, we’ll work through a self-contained, end-to-end data project together, comparing how each component might be handled in Python and SAS. Steps will include downloading data files from URLs, munging datasets together, summarizing data values, performing light data cleaning, building a basic statistical model, and saving results to external files. Along the way, we’ll also give a beginner-friendly overview of Python syntax and data structures, as well as important “gotchas” for SAS programmers.

This workshop is aimed at SAS programmers of all skill levels, including those with no prior experience using Python. A Google account will be needed to interact with code examples through Colab (https://colab.research.google.com/). All class materials, including complete setup instructions, will be made available through https://github.com/saspy-bffs/wuss-2024-python-how

192 : Utilizing the Medicare Part D Prescriber Public Use Data File as a “Training” Data Set
Louise S Hadden
Thursday, September 05, 2024, 10:00 AM – 11:00 AM, Location: Golden State

The Centers for Medicare & Medicaid Services (CMS) provides a public use dataset, the Part D Prescriber Public Use File which contains information on prescription drug events (PDEs) for Medicare beneficiaries with a Part D prescription drug plan, including information on prescribers, drug names, drug utilization, and drug costs. Although there is significant redaction in the data set to protect from beneficiary identification, the Part D Prescriber PUF can be a helpful tool in planning analyses using identified data, and in constructing analysis plans using Medicare PDE data and/or private drug utilization data. This Hands On Workshop (HOW) demonstrates the process of acquiring and reconstituting the Part D Prescriber PUF, exploring helpful additional data sources such as NDC crosswalks and morphine milligram equivalent (MME) calculations. Preliminary analyses and exercises will be conducted, delving into spending and utilization data, as well as prescriber information, brand name and generic name, total number of prescriptions for each drug, 30-day standardized fill counts, total days’ supply, and total drug costs. Using geographic information for each prescriber, prescriber locations will be geocoded and mapped using SAS PROC GEOCODE and SAS PROC SGMAP. Demonstrations will be conducted using SAS; however, participants are welcome to use the provided materials and exercises with the software package of their choice, such as R or Python.

198 : Commit early, commit often! A gentle, hands-on introduction to the joy of Git and GitHub
Isaiah Lankham, Matthew T Slaughter
Friday, September 06, 2024, 10:30 AM – 12:00 PM, Location: Golden State
Download Slide Presentation (PDF)

In this hands-on workshop, we’ll introduce you to the joy of Git and GitHub for managing codebases of any size, whether working alone or as part of a team.

Collaborating together, we’ll practice using the GitHub website. Topics will include basic Git/GitHub concepts like forking, cloning, and branching, as well as best practices for maintaining a well-organized history of code changes. We’ll also use the GitHub web interface for pull requests, which are the standard mechanism for contributing to open-source projects, and we’ll ensure every participant leaves this workshop with (a) a fully setup GitHub account and (b) at least one open-source contribution.

The social coding platform GitHub is synonymous with open-source software development, with many developers publishing their code as a form of résumé. Behind the scenes, GitHub uses software called Git, which was developed as a distributed version control system for managing contributions by thousands of developers to the Linux kernel.

No knowledge of Git or GitHub will be assumed, and no software will need to be installed. In order to work through interactive examples, accounts will be needed for GitHub and Google. Complete setup steps will be provided at https://github.com/saspy-bffs/wuss-2024-git-how

200 : Map It Out: Using SG Attribute Maps for Precise Control of PROC SGPLOT Output
Josh Horstman
Wednesday, September 04, 2024, 03:00 PM – 04:00 PM, Location: Golden State
Download White Paper (PDF)

The SGPLOT procedure, part of the ODS Statistical Graphics package, allows for extensive customization of nearly all aspects of plot output. These capabilities are commonly used to distinguish between groups or categories being compared through the use of distinct plot attributes, such as symbols and colors. However, there are times when it is advantageous to be able to associate specific plot attributes with specific data values. SG attribute maps provide functionality that does exactly that. This presentation will provide an introduction to the use of SG attribute maps in conjunction with PROC SGPLOT. A series of examples will demonstrate how attribute maps are used and why they are useful as a programming tool. Both discrete and range attribute maps will be used to modify a variety of plot attributes, such as plot marker symbols and colors, line styles and fill patterns.

207 : Mastering Clinical Trial Reporting with SAS
Matt Becker
Wednesday, September 04, 2024, 05:00 PM – 06:00 PM, Location: Golden State
Download Slide Presentation (PDF)

This hands-on workshop offers an immersive experience in using SAS software to generate essential components of clinical trial reporting, including data sets, tables, listings, and figures (TLFs). Participants will learn how to leverage SAS for data manipulation to CDISC standards , statistical analysis, and reporting, crucial for clinical trial documentation and regulatory submissions. The workshop will cover key topics such as data transformation, coding and validation of clinical trial data, and the creation of standard TLFs following industry guidelines. Attendees will gain practical skills in using SAS procedures and macros to automate and streamline the production process, ensuring accuracy and compliance with regulatory requirements. Through interactive sessions and real-world examples, participants will develop a deep understanding of how to effectively utilize SAS to support the entire lifecycle of clinical trials, from data preparation to the final presentation of results. This workshop is designed for clinical data managers, statisticians, and programmers seeking to enhance their proficiency in SAS and improve the efficiency and quality of their clinical trial reporting.

209 : GIT Intro for SAS Coders that need to GIT started
Zeke Torres
Friday, September 06, 2024, 08:30 AM – 10:00 AM, Location: Golden State
Download White Paper (PDF)

With the advent of more team diversity in skills syntax, the challenges have only increased in how to integrate new team members, code styles, code upgrades and syntax. Even if you are a single person creating and editing your own code, its time to Git started. This is the how to start if you don’t know how to Git started, especially if you are working with a team.

The need exists for teams to have a practical outline of elements incorporated to achieve a strategic harmony of SAS, Python, R (and more) with Git or some version control. Improve code transparency version control with useful team code review methods. On top of all that, we must manage the data we use with our code. It’s important to document a projects key decision points during an analytics cycle. Doing so increases the accuracy of analytics results by reducing issues due to missing key code documentation.

There are key differences in the way analytics teams work versus how data engineering and ETL data preparation coders work. It’s not just a syntax difference; its also a code cultural difference. This presentation will outline how these cultures can co-exist and identify areas where they can collaborate. It will discuss the strengths they each contribute, while still taking advantage of the common elements of code governance, documentation and code version control.

Live Demonstrations

121 : Jedi SAS Tricks for DS2 Programmers
Mark Jordan
Thursday, September 05, 2024, 11:00 AM – 11:30 AM, Location: Resource Central
Download White Paper (PDF)

If you’ve heard about the SAS DS2 programming language, you probably know that it combines the process control of the DATA step with the simplicity and power of SQL. You may even know about DS2’s simple syntax for parallel processing and ability to handle most ANSI data types at full precision. But there are many lesser-known superpowers built into the DS2 language. In this presentation, you will learn how to:
• Defer array dimensioning until execution
• Apply a format to all the elements of an array
• Use FIRST. / LAST. logic without pre-sorting your input data
• Use matrix math to resolve complex problems
• Write results to a CSV file even though DS2 does not read or write text files (using ODS)
• Use parameter-driven threads for super-flexible & reusable parallel processing code

126 : Five Easy Ways to Make Your SAS Code Run Faster
Mark Jordan
Friday, September 06, 2024, 10:30 AM – 11:00 AM, Location: Resource Central
Download Slide Presentation (PDF)

Don’t you just hate waiting for your results after submitting a SAS program? I sure do! In this session, I’ll demonstrate five easy “thumb rules” that you can use to help you write faster-running SAS code. These techniques are valid for code executed in SAS 9 or the SAS Viya Compute Server. I’ll demonstrate head-to-head comparisons of the different techniques in SAS 9, using SAS datasets and Oracle tables as data sources. You’ll see how simple coding choices affect execution time, and learn which choices provide the best performance in the circumstances we commonly encounter every day.

128 : Beyond Macro – Data-Driven Programming with CAS in SAS Viya
Mark Jordan
Wednesday, September 04, 2024, 03:30 PM – 04:30 PM, Location: Resource Central
Download White Paper (PDF)

With the adoption of SAS Viya accelerating processing in CAS is becoming more common, and CAS speaks CASL. Seasoned SAS coders often use SAS macro to produce data-driven programs, automating tedious programming tasks. When working with CAS from the SAS Compute Server, it’s important to know where, when, and how the CASL that executes is generated. Whatever your experience level, the interactions between SAS code, CASL, and Macro can be intimidating. This presentation aims to demystify that process.

135 : Handling Missing Values in SAS 9 and SAS Viya
Melodie Rush
Thursday, September 05, 2024, 04:30 PM – 05:30 PM, Location: Resource Central
Download Slide Presentation (PDF)

What do you do when you have missing values in your data? In SAS we have many ways to manage missing values. In this session, we cover what missing values are, why and when missing values occur, and how to manage missing values. We discuss functions, procedures, and how different products deal with missing values.

156 : Correlation of a Composite Self-Reported Symptoms Severity Score with Patient Number of Symptoms as Predictors of SARS-CoV-2 Viral Status/Load using the SAS(r) System
Jesse Albert Canchola
Thursday, September 05, 2024, 09:00 AM – 09:30 AM, Location: Resource Central
Download Slide Presentation (PDF)

In this prospective cohort study in Germany, one aim was to investigate the correlation of a symptoms severity score composed of patient self-reports with patient number of symptoms as predictors of SARS-CoV-2 status (positive/negative). Establishing a stronger link between the presence of virus and the severity of the disease could help in the early assessment of risk and predicting COVID-19 disease outcomes. It could also improve the determination of who is most in need of treatment and help reduce the spread of infection. A symptom severity score (S3) was constructed from the participant survey and validated (Cronbach’s alpha=0.7) then categorized into three disease severity score bins (S3C): asymptomatic, mild to moderate symptoms and severe symptoms. The S3 construct correlated with the total symptoms (Pearson r=0.963, p<0.0001). Furthermore, the categorized version of S3 correlated with the calculated number of symptoms categorized into three categories: no symptoms, 1-2 symptoms, and 3 or more symptoms (Spearman's r=0.988, p<0.0001). A generalized estimating equation (GEE) model using S3C as the predictor of SARS-CoV-2 status (positive/negative) showed that those participants who reported severe symptoms in the survey had an odds 6.5 times higher to be positive (vs. negative) than those who reported no symptoms (OR=6.5, 95% CI: 3.5 to 12.4, p<0.0001). Similar statistically significant results were found when comparing severe symptoms vs. mild to moderate symptoms (OR=2.3, CI: 1.3 to 4.1, p=0.0025) and mild to moderate symptoms vs. asymptomatic reports (OR=2.8, 95% CI: 1.4 to 5.4, p=0.0030). The SAS System v9.4 was used solely to produce the complete analysis.

163 : CTRL + S(HARE): A Programmer’s Advice for Branding on Social Media
Inka Leprince
Thursday, September 05, 2024, 10:30 AM – 11:00 AM, Location: Resource Central
Download White Paper (PDF)

Navigating social media as a programmer requires a blend of technical knowledge, personal branding, and community engagement. The foci of this presentation are the best social media practices and strategies for enhancing visibility through tagging individuals/companies, establishing regular posting schedules, and leveraging fun and interactive content. Resources for free design software will be provided as well as practical tips for integrating branding elements in the creation of visually striking graphics. By following these guidelines, you can effectively promote yourself and engage your followers.

164 : Here Comes the Rain (Cloud Plot) Again
Bart Jablonski
Wednesday, September 04, 2024, 05:00 PM – 05:30 PM, Location: Resource Central
Download White Paper (PDF)

Rain cloud plots are very popular and practical tools for data visualization. They enable the display of several components in one figure, including: kernel density estimates, data points, and box-and-whisker plots. This method facilitates the comparison of variable distribution across different categories. During this presentation, a macro called %RainCloudPlot() will be introduced and its capabilities for creating rain cloud plots will be presented. Sample code will accompany all examples.

165 : Streamlining Your Workflow: Creating Portable and Automated SAS Project Folders Using Your SAS Enterprise Guide Project Name
Chary Akmyradov
Thursday, September 05, 2024, 04:00 PM – 04:30 PM, Location: Resource Central
Download Slide Presentation (PDF)

In this presentation, I will demonstrate how to leverage the SAS Enterprise Guide (EG) project name and directory location to automate the definition of libraries and the creation of project-related folders. The session will cover the application of Autoexec, DLCreateDir option, the automatic macro variable &_ClientProjectPath, and essential SAS functions such as dequote, find, and substr.

Attendees will learn how to:

1. Define a library using a portion of the SAS EG project name.
2. Assign an automatically created (if not existing) library folder.
3. Create additional folders for exports, reports, and more.

A key feature of this approach is that when the SAS EG project is executed from a new directory, the library location is updated automatically, ensuring the project’s portability. This session will provide practical insights and step-by-step instructions to enhance your SAS project management and workflow efficiency. This demonstration is suitable for all levels of SAS programmers.

168 : PROC BGLIMM: The Smooth Transition to Bayesian Analysis
Danny R Modlin
Friday, September 06, 2024, 11:00 AM – 11:30 AM, Location: Resource Central
Download Slide Presentation (PDF)

Many analysts are interested in taking models they currently have and transitioning them to the Bayesian realm. Most leap from their favorite classical analysis procedure directly to PROC MCMC, the general-purpose Bayesian procedure. This presentation will feature the BGLIMM procedure available since SAS/STAT 15.1. This will allow the participant to model non-normal responses and include random effects within their Bayesian approach. Discussion will include options of priors and availability of statements. Examples will include models originally written in PROCs REG, GLM, GLMSELECT, GENMOD, MIXED, and GLIMMIX.

173 : SAS Visual Analytics a low code no code approach to dashboard reporting
Charu Shankar
Thursday, September 05, 2024, 02:00 PM – 03:00 PM, Location: Resource Central
Download Slide Presentation (PDF)

Are you a SAS programmer with a creative streak, frustrated by the hours spent rewriting customized reporting procedures? Wish you had an easy, visual option to streamline your work? Want to learn hacks to cut down coding time and focus more on designing reports? Wishing you could quickly pick up reporting best practices that you can implement right away with a low-code, no-code approach? Curious about the distinction between traditional reporting and dashboard reporting? Wondering what adopting SAS Visual Analytics might look like & how it might enhance your reporting process? ? If you’re nodding along to any of the above, then this session is for you.

In this session, we will cover
• How to apply the principles of good design in reporting
• How to utilize the main report objects and their practical use
• How to dynamically visualize Cirque du Soleil events using a SAS Visual Analytics Dashboard

All levels welcome!

178 : The Dataset-JSON Submission Process
Sarvar Khamidov, James Joseph
Wednesday, September 04, 2024, 05:30 PM – 06:00 PM, Location: Resource Central
Download Slide Presentation (PDF)

The submission of clinical trial data to regulatory bodies traditionally has involved complex, time-consuming processes reliant on old SAS V5 XPT format.

Through a detailed comparison, this study highlights the key differences in data submission workflows, focusing on aspects such as data preparation, validation, submission, and review processes. The paper also discusses the implications of adopting JSON datasets in terms of interoperability, regulatory acceptance, and the impact on clinical trial timelines.

Empirical data from recent clinical trials is utilized to demonstrate the practical benefits and potential limitations of transitioning to JSON-based submissions. Our findings suggest that while the XPT format has served the industry well, the JSON format offers significant advantages in terms of flexibility and automation, which could lead to more efficient and faster submissions.

Ultimately, this paper aims to provide insights for researchers, data managers, and regulatory authorities considering the shift from XPT to JSON, offering a roadmap for navigating this transition and leveraging the full potential of modern data submission technologies.

184 : Integration of an R Language Function Within SAS to Perform Conditional Multivariate Saddlepoint Approximation
Christopher Johnson
Thursday, September 05, 2024, 03:00 PM – 03:30 PM, Location: Resource Central
Download White Paper (PDF)
Download Slide Presentation (PDF)

Issues with small and sparse data are well known to produce bias in maximum likelihood estimates (MLEs) when performing logistic regression analyses. A common alternative to ordinary maximum likelihood based logistic regression is exact logistic regression, whereby an exact distribution of sufficient statistics for one or more parameters of interest is generated, conditional upon the sufficient statistics for one or more nuisance parameters. Through this generated exact conditional distribution, conditional maximum likelihood estimation may be performed to obtain accurate parameter estimates and conduct hypothesis testing. A drawback of this approach is that parameters are estimated fully conditionally when more than one parameter of interest exists. That is, each parameter of interest is estimated separately and completely conditional on all other parameters of interest. As a result, parameter interpretation in the context of other parameters is not equal to that of common maximum likelihood estimation-based logistic regression procedures in which parameter estimation and hypothesis testing is performed jointly. An alternative to CMLE is an extremely accurate and efficient higher-order asymptotic method known as the saddlepoint approximation, which may be used to produce joint parameter estimates with interpretation akin to those produced via MLE. Here, we provide a general overview of a novel multivariate saddlepoint approximation and feature its implementation in a SAS macro with an embedded R function. Using this macro, multivariate saddlepoint approximations are obtained to provide a robust alternative analysis to small and sparse data scenarios when ordinary logistic regression is not recommended.

201 : How RANK are your deciles? Using PROC RANK and PROC MEANS to create deciles based on observations and numeric values
Lisa A Mendez
Thursday, September 05, 2024, 08:30 AM – 09:00 AM, Location: Resource Central
Download White Paper (PDF)
Download Slide Presentation (PDF)

For many cases using PROC RANK to create deciles works sufficiently, but occasionally, you find that it
does not work for your needs. PROC RANK uses number of observations to produce a rank; however, if
you need weighted percentiles then PROC RANK will not work. Instead, you can use Proc Means to
successfully create weighted percent groups. This paper will illustrate the basic usage of PROC RANK
and how to use PROC MEANS for the alternative. The paper will utilize BASE SAS® 9.4 code and will
use a fictional dataset that provides the total number of prescriptions written by providers for two years.
All levels of SAS users may benefit from the information provided in this paper.

206 : SAS Viya in Action: Enhancing Life Sciences Data Analysis
Matt Becker
Thursday, September 05, 2024, 03:30 PM – 04:00 PM, Location: Resource Central
Download Slide Presentation (PDF)

The demonstration of SAS Viya for Life Sciences showcases the platform’s capabilities in transforming complex data analytics and modeling into actionable insights. SAS Viya offers robust tools for data management, advanced analytics, and artificial intelligence (AI) tailored to the unique needs of the life sciences sector. This demo highlights key functionalities including data integration from diverse sources, real-time data processing, and advanced analytics techniques such as machine learning and predictive modeling. The platform facilitates seamless collaboration across multidisciplinary teams, enhances data transparency and traceability, and ensures compliance with regulatory standards. By leveraging SAS Viya, life science organizations can accelerate drug discovery, optimize clinical trial processes, and improve patient outcomes through precise, data-driven decision-making. The demonstration underscores the platform’s scalability, flexibility, and user-friendly interface, making it an essential tool for driving innovation and efficiency in life sciences research and development.

211 : SAS Studio, SAS Enterprise Guide, SAS Extension for Visual Studio Code: Which should I use?
Casey Smith
Friday, September 06, 2024, 08:30 AM – 09:30 AM, Location: Resource Central
Download Slide Presentation (PDF)

SAS offers several client applications (namely, SAS Studio, SAS Enterprise Guide, SAS Extension for Visual Studio Code, and SAS Display Management System) for SAS programming and other features, such as flow building, data preparation, analysis, and ad-hoc querying and reporting. It can be a little overwhelming to know which SAS client to use! In this demo, I will show, compare, and contrast the features of each of these client applications, highlight the pros and cons of each, and give you a better idea of which one would fit your needs best.

212 : Using VS Code with SAS
Chris Hemedinger
Wednesday, September 04, 2024, 03:00 PM – 03:30 PM, Location: Resource Central
Download Slide Presentation (PDF)

Discover the power of SAS programming within Visual Studio Code with the official SAS extension. This session will delve into the extension’s capabilities, including seamless integration for SAS developers, support for SAS 9.4 and SAS Viya, and flexible connection options via access tokens, SSH, and IOM. Learn unique tips and tricks to enhance your productivity and get an overview of the open source project and support model that underpins this innovative tool.

This presentation is a must-attend for those looking to streamline their SAS programming workflow and engage with the vibrant community supporting the SAS extension for VS Code.

213 : Working with SAS and Microsoft 365
Chris Hemedinger
Thursday, September 05, 2024, 09:30 AM – 10:00 AM, Location: Resource Central
Download Slide Presentation (PDF)

Modernize your approach to accessing and publishing traditional Microsoft Office content with SAS. Come see how you can unlock new power and flexibility by using the Microsoft Graph APIs from your SAS programs. You will learn how to: Connect to Microsoft 365 from within your SAS program, explore your OneDrive and SharePoint/Teams files, and download and upload content into your online cloud-hosted folders – all with a code-centric approach. You will also learn about a library of SAS macro routines that help to make the most common tasks easy to accomplish.

216 : Who’s Still Bringing That Big Data Energy? Leading a SAS® Conference by Leveraging a 48-Year Longitudinal Analysis of 30,498 Presentations in the SAS User Community
Troy Martin Hughes
Thursday, September 05, 2024, 11:30 AM – 12:00 PM, Location: Resource Central
Download White Paper (PDF)

This analysis examines presentations at SAS® user group conferences between 1976 and 2024, and supplements the introductory analysis that the author debuted at WUSS 2023. (Hughes, 2023) Both analyses evaluate white papers maintained on www.LexJansen.com (aka “the LEX”), as well as other presentations referenced on this inimitable site. Presentations are drawn from multiple conferences, including: SAS User Group International (SUGI, may she rest in peace), SAS Global Forum (SGF, may she be revived), SAS Explore, Western Users of SAS Software (WUSS), Midwest SAS Users Group (MWSUG), South Central SAS Users Group (SCSUG), Southeast SAS Users Group (SESUG), Northeast SAS Users Group (NESUG), Pacific Northwest SAS Users Group (PNWSUG), and Pharmaceutical Software Users Group (PharmaSUG). This current analysis evaluates the critical role that the 2023 analysis played in data-driven decision-making for the WUSS 2024 conference. As always, unremittent thanks to a friend, mentor, and fellow tall guy, “the” Lex Jansen.

217 : From Word Clouds to Phrase Clouds to Amaze Clouds: A Data-Driven Python Programming Solution To Building Configurable Taxonomies That Standardize, Categorize, and Visualize Phrase Frequency
Troy Martin Hughes
Friday, September 06, 2024, 09:30 AM – 10:00 AM, Location: Resource Central
Download White Paper (PDF)

Word clouds visualize the relative frequencies of words in some body of text, such as a website, white paper, blog, or book. They are useful in identifying contextual focus and keywords; however, word clouds—as commonly defined and implemented—suffer numerous limitations. First, multi-word phrases such as “data set” or “Base SAS” are unfortunately segmented into single words—”data,” “set,” “Base,” and “SAS.” Second, desired capitalization often cannot be specified, such as visualizing “PROC PRINT” even when its lowercase “proc print” is observed in text or code. Third, spelling variations (e.g., single and plural nouns, various verb conjugations, abbreviations and acronyms) are not mapped to each other. Similarly, and fourth, comparable words or phrases (e.g., “PROC PRINT” and “PRINT procedure”) are not mapped to each other, representing a further lack of entity resolution. This text and its Python Pandas solution seek to overcome data quality, data integrity, and data standardization issues that plague word clouds, by defining and applying configurable taxonomies—data models that can impart more meaning and precision to ultimate word/phrase cloud visualizations. The result is a phrase cloud that amazes—an amaze cloud!

218 : An introduction to Quarto: A Versatile Open-Source Tool for Data Reporting and Visualization
Joshua J Cook, Swann Arp Adams
Friday, September 06, 2024, 11:30 AM – 12:00 PM, Location: Resource Central
Download White Paper (PDF)

In the collaborative landscape of data analysis, a common frustration among analysts stems from the need to integrate and harmonize different programming languages within a team. Teams often comprise interdisciplinary researchers, each with their unique programming preferences and expertise, leading to complexities in project integration and continuity. The difficulty in compiling and executing data projects cohesively can hinder efficiency and impede the delivery of coherent, multi-faceted data insights. This challenge necessitates a solution that can bridge the gaps between varying coding languages and methodologies to streamline team collaboration and project completion.

Addressing this issue, the point of this paper is to present Quarto as an innovative solution that can unify the diverse programming approaches within a team. Quarto stands out by offering extensive cross-language support, enabling the integration of code from multiple languages into a singular, dynamic report. This versatile reporting system is tailored for the pharmaceutical and biotech industries, facilitating the creation of comprehensive reports and visualizations that cater to stakeholders at all technical levels. With Quarto, consolidating code, narrative text, and outputs into one document is seamless, accommodating outputs in various formats such as HTML, PDF, Word, Typeset, Markdown, PowerPoint, dashboards, websites, manuscripts, and even entire books. This paper serves as an introduction to Quarto’s capabilities, highlighting its role in enhancing collaboration and efficiency in data science projects across the spectrum of technical expertise.

219 : SAS-tastic Visuals: Turning Data into Eye Candy!
Matt Becker
Friday, September 06, 2024, 10:00 AM – 10:30 AM, Location: Resource Central
Download Slide Presentation (PDF)

In today’s data-driven world, the ability to visualize data effectively is crucial for making informed decisions. This presentation will dive into the powerful capabilities of SAS Graphics, showcasing how it can transform complex data into compelling visual insights. We will explore key features, including the creation of simple graphs, customized plots, and integration of advanced statistical visualizations. Whether you are a beginner or an experienced SAS user, this session will equip you with practical tips and techniques to elevate your data visualization skills, making your reports not just informative but visually engaging. Join us to discover how SAS Graphics can help you communicate your data story more effectively.

Lively Lectures

102 : Just Stringing Along: FIND Your Way to Great User-Defined Functions
Richann Jean Watson, Louise S Hadden
Thursday, September 05, 2024, 03:00 PM – 03:30 PM, Location: Regency E
Download White Paper (PDF)

SAS® provides a vast number of functions and subroutines (sometimes referred to as CALL routines). These useful scripts are an integral part of the programmer’s toolbox, regardless of the programming language. Sometimes, however, pre-written functions are not a perfect match for what needs to be done, or for the platform that required work is being performed upon. Luckily, SAS has provided a solution in the form of the FCMP procedure, which allows SAS practitioners to design and execute User-Defined Functions (UDFs). This paper presents two case studies for which the character or string functions SAS provides were insufficient for work requirements and goals and demonstrate the design process for custom functions and how to achieve the desired results.

103 : Incorporating Macro with PROC REPORT Code
Jane Eslinger
Friday, September 06, 2024, 08:30 AM – 09:30 AM, Location: Regency E
Download White Paper (PDF)

PROC REPORT is used across many industries to generate reports that often need to be generated on a weekly, monthly, or quarterly basis. The PROC REPORT code for these reports must be robust enough to handle things like new data, different dates, and changes to titles and headers. Busy programmers don’t have time to update the program for every new run, but smart programmers know that SAS Macro must be utilized. This paper explores how to incorporate SAS Macro into PROC REPORT code, going beyond simple macro variable references. Through examples, you will learn how to write code that both iterates over a series of macro variables and parses a singular macro variable. Both use cases can be utilized by either placing PROC REPORT inside a macro program and employing macro do group logic to generate multiple statements like DEFINE statements or compute blocks, or by calling a separate macro program from within the PROC REPORT code which allows the macro program to be shared across programs for report standardization.

104 : SAS® Macro Programming Tips and Techniques
Kirk Paul Lafler
Thursday, September 05, 2024, 02:00 PM – 03:00 PM, Location: Regency C
Download White Paper (PDF)

The SAS® Macro Language is a powerful feature for extending the capabilities of the SAS System. This paper highlights a collection of techniques for constructing reusable and effective macro tools. Attendees are introduced to the techniques associated with building functional macros that process statements containing SAS code; design reusable macro techniques; create macros containing keyword and positional parameters; utilize defensive programming tactics and techniques; build a library of macro utilities; interface the macro language with the SQL procedure; and develop efficient and portable macro language code.

108 : Creating Custom Excel Spreadsheets with Built-in Autofilters Using SAS® Output Delivery System (ODS)
Kirk Paul Lafler
Friday, September 06, 2024, 11:00 AM – 11:30 AM, Location: Regency E
Download White Paper (PDF)

Spreadsheets have become the most popular and successful data tool ever conceived. Current estimates show that there are more than 750 million Excel users worldwide. A spreadsheet’s simplicity and ease of use are two reasons for the growth and widespread use of Excel around the globe. Additional value-added features have also helped to expand the spreadsheet usefulness among a growing number of users including its collaborative capabilities, being customizable, ability to manipulate data, application of data visualization techniques, mobile device usage, automation of repetitive tasks, integration with other software, data analysis, and filtering capabilities using autofilters. This last value-added feature, filtering with autofilters, is the theme for this paper. An example application will be illustrated that creates a custom Excel spreadsheet with built-in autofilters, or filters that provide users with the ability to make choices from a list of text, numeric, or date values to find data of interest quickly, using the SAS® Output Delivery System (ODS) Excel destination and the REPORT procedure.

112 : Introduction to developing Neural Nets
David Horvath
Thursday, September 05, 2024, 08:30 AM – 09:30 AM, Location: Regency C

Neural nets have been discussed in in SAS regional and international conferences since last century, This session will be a bit different in that it starts with the obligatory basic background and then goes into actual code. We will discuss what Machine Learning can and cannot solve. There will be more discussion of the how results are produced and less about the overall concepts.

Topics
* What is a neuron?
* How does a neural net model a human neuron?
* How does that work with data?
* What does it NOT solve? Discuss the biggest problems with data.
* Examples
1. Working data
2. Why’d we get that result?
3. More Working data
4. Why’d we get *that* result?
5. Reviewing the code

113 : Leveraging “UNIX Tools” (GNU) for Data Analysis
David Horvath
Wednesday, September 04, 2024, 03:00 PM – 04:00 PM, Location: Regency A

Life would be so much easier if everything was in a database or readily processed within SAS®. But that
is not the case. All too often we get data files (or have to send them) in various formats. This session
discusses some of the tools available to help you figure out what the file looks like so you can pull it apart
using those tools or your SAS. While the GNU version of these tools will be the focus, the skills learned
apply to many different platforms (Microsoft’s Bash under Windows 10, Cygwin under Microsoft Windows,
MAC OSX, the Linux core of Android, commercial Linux — like Red Hat Enterprise, and commercial UNIX
— like IBM’s AIX or Sun/Oracle’s Solaris).
Of particular interest are ‘head’, ‘tail’, ‘wc’, ‘awk’, ‘dd conv’, and shells.
A few of the differences between UNIX/Linux and Windows will also be discussed in case you ever have
to deal with those environments in our heterogeneous environments. This knowledge also comes in
handy if you need to migrate code from an existing UNIX/Linux-based application.

114 : From Muggles to Macros: Transfiguring Your SAS® Programs With Dynamic, Data-Driven Wizardry
Josh Horstman, Richann Jean Watson
Wednesday, September 04, 2024, 05:00 PM – 06:00 PM, Location: Regency E
Download Slide Presentation (PDF)

The SAS macro facility is an amazing tool for creating dynamic, flexible, reusable programs that automatically adapt to change. This presentation uses examples to demonstrate how to transform static “muggle” code full of hardcodes and data dependencies by adding macro language magic to create data-driven programming logic. Cast a vanishing spell on data dependencies and let the macro facility write your SAS code for you!

115 : Developing Artificial and Convolutional Neural Networks with Python’s Keras API for TensorFlow
Ryan Paul Lafler, Anna T. K. Wade
Friday, September 06, 2024, 08:30 AM – 09:30 AM, Location: Regency D
Download White Paper (PDF)

Capable of accepting and mapping complex relationships hidden within structured and unstructured data, neural networks are built from layers of neurons and activation functions that interact, preserve, and exchange information between layers to develop highly flexible and robust predictive models. Neural networks are versatile in their applications to real-world problems; capable of regression, classification, and generating entirely new data from existing data sources, neural networks are accelerating recent breakthroughs in Deep Learning methodologies. Given the recent advancements in graphical processing unit (GPU) cards, cloud computing, and the availability of interpretable APIs like the Keras interface for TensorFlow, neural networks are rapidly moving from development to deployment in industries ranging from finance, healthcare, climatology, video streaming, business analytics, and marketing given their versatility in modeling complex problems using structured, semi-structured, and unstructured data. This paper explores fundamental concepts associated with neural networks including their inner workings, their differences from traditional machine learning algorithms, and their capabilities in supervised, unsupervised, and generative AI workflows. It also serves as an intuitive, example-oriented guide for developing Artificial Neural Network (ANN) and Convolutional Neural Network (CNN) architectures using Python’s Keras and TensorFlow libraries for regression and image classification tasks.

116 : Charting Your Organization’s Machine Learning Roadmap
Ryan Paul Lafler
Wednesday, September 04, 2024, 02:30 PM – 03:00 PM, Location: Regency A
Download White Paper (PDF)

Machine Learning is experiencing a golden age of investment, democratization, and accessibility across all sectors and industries encompassing the life sciences, healthcare, financial technology (fintech), consumer marketing, e-commerce, manufacturing, and more. But what exactly is “Machine Learning”? How is it connected to Artificial Intelligence (AI)? And most importantly, how can data scientists, programmers, software engineers, and/or researchers start their endeavors into Machine Learning? This presentation answers these questions, and more, by giving attendees a roadmap to help them navigate the complexities of Machine Learning in an application-oriented guide.

This presentation covers the main aspects of Machine Learning including supervised, unsupervised, and semi-supervised approaches as well as Deep Learning. Attendees are given a roadmap that starts with linear regression and progressively builds towards more complex and flexible algorithms with discussions about the advantages and disadvantages of using certain algorithms over others. In doing so, attendees will learn about Python libraries for Machine Learning; real-world applications of both labeled and unlabeled data; overfitting and underfitting; cross-validation; and the importance of hyperparameter tuning to better fit algorithms to their data.

117 : LAST CALL to Get Tipsy with SAS®: Tips for Using CALL Subroutines
Lisa A Mendez, Richann Jean Watson
Friday, September 06, 2024, 10:30 AM – 11:00 AM, Location: Regency E
Download White Paper (PDF)

This paper provides an overview of six SAS CALL subroutines that are frequently used by SAS®
programmers but are less well-known than SAS functions. The six CALL subroutines are CALL MISSING, CALL SYMPUTX, CALL SCAN, CALL SORTC/SORTN, CALL PRXCHANGE, and CALL EXECUTE. Instead of using multiple IF-THEN statements, the CALL MISSING subroutine can be used to quickly set multiple variables of various data types to missing. CALL SYMPUTX creates a macro variable that is either local or global in scope. CALL SCAN looks for the nth word in a string. CALL SORTC/SORTN is used to sort a list of values within a variable. CALL PRXCHANGE can redact text, and CALL EXECUTE lets SAS write your code based on the data. This paper will explain how those six CALL subroutines work in practice and how they can be used to improve your SAS programming skills.

118 : Can You Teach An Old Dog New Tricks?
Scott Burroughs
Wednesday, September 04, 2024, 04:00 PM – 04:30 PM, Location: Golden State
Download White Paper (PDF)
Download Slide Presentation (PDF)

It’s a fact of life that things change, and that also applies to our jobs, regardless of industry. With our jobs being technical, things will likely change even faster. It may or may not seem surprising that SAS is still the main software used in this industry, but when a ‘free’ alternative comes out, companies are going to consider it. Additionally, our jobs may change (to a new company or within our companies) and so software platforms are likely to change. As a self-professed old-timer, can I learn these new ways of doing things after having done them other ways (maybe even one way) for so long?

119 : Limit of Detection Calculation Methods Comparison for PCR-based Quantitative Studies using the SAS® System
Jesse Albert Canchola
Friday, September 06, 2024, 10:30 AM – 11:00 AM, Location: Regency C
Download White Paper (PDF)
Download Slide Presentation (PDF)

In assay performance evaluation for quantitative assay studies based on polymerase chain reaction (PCR), the Limit of Detection (LoD) is defined as the lowest concentration or amount of analyte that is consistently detectable (in our case, in at least 95% of the samples tested; CLSI EP17-A2). In practice, the estimation of the LoD uses a parametric curve fit to a set of panel member (PM1, PM2, PM3, etc.) data where the responses are binary (i.e., percent detected at a certain level). Typically, the parametric curve fit to the percent detection levels takes on the form of a logistic or probit distribution. Additional methods for estimation of the LoD in PCR-based studies include using the “based on hit rate” and maximum likelihood estimation (MLE) methods (Singh & Nocerino, 2001). Amongst all methods, the MLE method is preferred since the MLE is taken to be sufficient (Rice, 1988) given the selected parent probability distribution assumed to model the data. Moreover, it can be shown that the MLE for a Poisson-distributed random variable has the minimum variance unbiased estimator (MVUE) property (Cassella and Berger, 2002), thus allowing it to be used as a reference in comparison with other methods. We performed a comparison of the LoD calculation methods using various realizations of a real seven-member panel for HIV-1 PCR assay results.

120 : SAS Job Searching and Interviewing tips – Strategies in the Post-Pandemic era
Jayanth Iyengar
Wednesday, September 04, 2024, 02:30 PM – 03:00 PM, Location: Regency D
Download White Paper (PDF)
Download Slide Presentation (PDF)

Searching for work in the data analytics market is more competitive than ever before. Part of this is due
to the nature of the work environment, which has shifted from in office, on-site to remote and hybrid
positions. Also, because of the emergence and growth of open-source tools, such as R and
Python, candidates for SAS positions need to be familiar with these coding tools. In this paper, I’ll outline
tips and strategies for success in the application and interview process in the post-pandemic era. I’ll also
discuss the interview process in detail, from initial HR interviews to final interviews with a hiring
manager.

122 : Airbnb Consumer Behavior Analysis
Evelyn Larson, Anabeli Chandra
Friday, September 06, 2024, 09:30 AM – 10:00 AM, Location: Regency D
Download White Paper (PDF)

“Oh, the places you’ll go.” In the dynamic landscape of business, understanding the customer is key. This research delves into the intricate tapestry of Airbnb user trends to predict whether a customer will book their stay with Airbnb after creating an account. We then want to assess whether we can determine which country a customer will book. The primary objective is to ensure customer retention and assess demand, enabling predictive modeling, optimized marketing strategies, and enhanced customer satisfaction. The methodology of our research included a multifaceted approach to extract meaningful patterns from the vast dataset. JMP was used to partition the data into a manageable subset useful in analysis. We then utilized JMP to preprocess and cleanse the data, ensuring its reliability. Predictive analysis via SAS Model Studio, JMP, and Python was used to determine where customers will book. Subsequently, clustering techniques were applied to segment Airbnb users into distinct categories based on shared characteristics. Data visualization was then conducted using JMP to provide a comprehensive overview of the identified consumer segments. Further analysis was conducted to refine the insights and validate the robustness of our findings. Preliminary results showcase distinct consumer segments within the Airbnb user base, each characterized by unique preferences and behaviors, and determined where similar customers will book in the future to help anticipate demand. Python, clustering, and JMP collectively contribute to a nuanced understanding of Airbnb user segmentation. This research is expected to have far-reaching implications for growing hospitality businesses seeking to tailor their strategies to diverse consumer segments. The refined categorization of Airbnb users allows for targeted marketing, personalized offerings, more effective business decisions, demand predictions, and increased profitability.

123 : How does SAS Viya and Open Source Integrate
Melodie Rush
Wednesday, September 04, 2024, 02:30 PM – 03:30 PM, Location: Regency E
Download Slide Presentation (PDF)

Combining the power of SAS with open source technologies allows you to unify disparate tools and analytic assets into a streamlined, collaborative environment that fosters productivity, business agility, and tangible results. In this session designed to enhance your understanding of this integration, you’ll explore the many ways SAS® Viya® works seamlessly with open source. Participants will discover how to leverage their existing SAS or other programming skills to access and harness the capabilities of SAS Viya. The session will cover practical applications such as using Python or R in the analytical flow of pipelines, as well as utilizing these languages through the SWAT (SAS Wrapper for Analytics Transfer) package. Additionally, attendees will learn how to use the Python Editor within SAS Studio, further enhancing their ability to perform sophisticated analyses and streamline their data science projects. This training is an excellent opportunity for users to maximize their use of SAS tools in conjunction with open source programming to achieve more efficient and impactful analytical outcomes.

124 : Macro Variable Arrays Made Easy with macroArray SAS package
Bart Jablonski
Friday, September 06, 2024, 11:00 AM – 12:00 PM, Location: Regency C
Download White Paper (PDF)

A macro variable array is a jargon term for a list of macro variables with a common prefix and numerical suffixes. Macro arrays are valued by advanced SAS programmers and often used as “driving” lists, allowing sequential metadata for complex or iterative programs. Use of macro arrays requires advanced macro programming techniques based on indirect reference (aka, using multiple ampersands &&), which may intimidate less experienced programmers. The aim of the paper is to introduce the macroArray SAS package. The package facilitates a solution that makes creation and work with macro arrays much easier. It also provides a “DATA-step-arrays-like” interface that allows use of macro arrays without complications that arise from indirect referencing. Also, the concept of a macro dictionary is presented, and all concepts are demonstrated through use cases and examples.

125 : Fifty Shades of SAS Programming, or 53 (+3) Syntax Snippets for a Table Look-up Task, or How to Learn SAS by Solving Only One Exercise!
Bart Jablonski, Quentin McMullen
Thursday, September 05, 2024, 09:30 AM – 10:30 AM, Location: Regency E
Download White Paper (PDF)

It is said that a good programmer should be lazy. And what about a good programming teacher?
We dare to say the same is true. This article will show that you can be lazy and also be able to teach an entire SAS course at the same time!

The aim of the article is to present a variety of examples of how to do one of the most common data processing programming tasks, table look-up, in SAS. We don’t assess these methods from a benchmarking or performance perspective but rather present them as an intellectual puzzle. Our goal is to explore how much SAS syntax (statements, PROCs, functions, etc.) could be taught using only one exercise. If you are a fan of unorthodox SAS programming, curious to learn about the variety and flexibility of the SAS language, or an innovative SAS teacher – this presentation is for you!

130 : Matrices and More: An Overview of the SAS/IML Language
Rick Wicklin
Wednesday, September 04, 2024, 05:00 PM – 06:00 PM, Location: Regency D
Download Slide Presentation (PDF)

Do you need a statistic that is not computed by any SAS® procedure? This talk introduces the SAS/IML® language to statistical programmers. The talks shows how to create and manipulate matrices, write user-defined functions, and implement custom analyses.

131 : Ten tips for simulating data with SAS
Rick Wicklin
Thursday, September 05, 2024, 11:00 AM – 12:00 PM, Location: Regency E
Download Slide Presentation (PDF)

Data simulation is a fundamental tool for statistical programmers. This paper presents 10 techniques that enable you to write efficient simulations in SAS. Examples include how to simulate from a probability distribution and how to approximate the sampling distribution of a statistic.

132 : Automation of Submission of Programs in Analysis Data Reviewer’s Guide for FDA and PMDA Submissions
Junying Zhang
Thursday, September 05, 2024, 03:00 PM – 03:30 PM, Location: Regency D
Download White Paper (PDF)

Analysis Data Reviewer’s Guide (ADRG) is recommended as an important part of a standards-compliant analysis data submission for clinical trials. ADRG Section 7 itemizes the programs that are included in the submission. This paper presents three useful SAS macros to automatically create Section 7 Submission of Programs for ADRG. MACRO1 automatically creates 7.1 ADaM Programs, MACRO2 creates 7.2 Analysis Output Programs and MACRO3 creates 7.3 Macro Programs. Linux command, Perl Regular Expression and SAS functions are used to extract the related information. The three macros have an option to create the tables according to FDA and PMDA requirements and add common customization options per different study requirements.

Section 7.1 ADaM Programs lists the ADaM program name, output, and macros used in programs. MACRO1 uses SAS PRXMATCH function to find ADaM programs. And SAS PRXCHANGE and SCAN function are used to search ADaM program logs to extract the input ADaM or SDTM datasets.

Section 7.2 Analysis Output Programs displays the Program Name, Output, Title and Input Datasets of the tables/figures programs submitted. Sometimes, there will be more than 100 table and figure outputs, and it will take the programmers a lot of time if they manually fill in this information. MACRO2 will automate Section 7.2.

Section 7.3 Macro Programs lists all macro programs called by the ADaM or tables/figures programs submitted. MACRO3 first uses PIPE and LS command to get macros used in ADAM or tables/figures programs. For FDA submission, we only display the local macros. For PMDA submission, we display both local and Gilead global macros.

The three macros automatically create the three tables in ADRG Section 7 and save a lot of time for programmers. And the customization options also satisfy most common study requirements. Other customization options could be easily added based on our macros.

133 : A Survey of Some of the Most Useful SAS Functions
Ron Cody
Wednesday, September 04, 2024, 03:00 PM – 04:00 PM, Location: Regency D
Download Slide Presentation (PDF)

SAS Functions provide amazing power to your DATA step programming. Some of these functions are essential—some of them save you writing volumes of unnecessary code. This talk covers some of the most useful SAS functions. Some of these functions may be new to you and they will change the way you program and approach common programming tasks.
The majority of the functions described in this talk work with character data. There are functions that search for strings, others that can find and replace strings or join strings together. Still others that can measure the spelling distance between two strings (useful for “fuzzy” matching). Some of the newest and most amazing functions are not functions at all, but call routines. Did you know that you can sort values within an observation? Did you know that not only can you identify the largest or smallest value in a list of variables, but you can identify the second or third or nth largest of smallest value? A knowledge of the functions described here will make you a much better SAS programmer.

136 : Navigating Career Crossroads: From Statistical Manager to Entrepreneur
Iuliana Constantin
Thursday, September 05, 2024, 10:00 AM – 10:30 AM, Location: Regency C
Download White Paper (PDF)

Losing my position as Statistical Manager in November 2023 due to a restructuring process was a moment of reflection and transformation. Although the decision was unrelated to my professional performance and driven by cost-cutting measures, it prompted me to think critically about my career and what I really wanted to do. Since I have been working as a Statistical Programmer for almost 16 years, I enjoyed analyzing data, and working with my team on the current studies, but was the time for me to do something else? This paper is about my personal experience of being laid off and embracing entrepreneurship. Throughout the narrative, I offer insights and strategies for navigating the transition from employee to small business owner, considering both the positive and negative effects of being an entrepreneur. I hope throughout this white paper, I can offer useful lessons for others considering a similar leap.

143 : The Essentials of SAS® Dates and Times
Derek Morgan
Thursday, September 05, 2024, 02:00 PM – 03:00 PM, Location: Regency E
Download White Paper (PDF)

The first thing you need to know is that SAS® software stores dates and times as numbers. However, this isn’t the only thing you need to know. This presentation gives you a solid base for working with dates and times in SAS. It introduces you to functions and features that enable you to manipulate your dates and times with surprising flexibility. This paper shows you some of the possible pitfalls with dates (and times and datetimes) in your SAS code and how to avoid them. We show you how SAS handles dates and times through examples, including the ISO 8601 formats and informats, and how to use dates and times in TITLE and FOOTNOTE statements. The paper closes with a brief discussion of Excel conversions.

144 : PROC SORT (then and) NOW
Derek Morgan
Friday, September 06, 2024, 10:00 AM – 10:30 AM, Location: Regency E
Download White Paper (PDF)

The SORT procedure has been an integral part of SAS® since its creation. The sort-in-place paradigm made the most of the limited resources at the time, and almost every SAS program had at least one PROC SORT in it. The biggest options at the time were to use something other than the IBM procedure SYNCSORT as the sorting algorithm, or whether you were sorting ASCII data versus EBCDIC data. These days, PROC SORT has fallen out of favor; after all, PROC SQL enables merging without using PROC SORT first, while the performance advantages of HASH sorting cannot be overstated. This leads to the question: Is the SORT procedure still relevant to any other than the SAS novice or the terminally stubborn who refuse to HASH? The answer is a surprisingly clear “yes”. PROC SORT has been enhanced to accommodate twenty-first century needs, and this paper discusses those enhancements.

145 : Time Since Last Dose: Anatomy of a SQL Query
Derek Morgan
Friday, September 06, 2024, 11:30 AM – 12:00 PM, Location: Regency E
Download White Paper (PDF)

Even though much of what we need to accomplish as SAS programmers can be accomplished using the DATA step, SQL can provide an alternative to large amounts of data manipulation. First, this paper will walk you through how SQL joins two datasets. Second, we will walk through the case of determining time since last (or first) dose and present an SQL solution in place of multiple SORTs, MERGEs, TRANSPOSEs, and use of the LAG function. Not only is the code economical, but the execution is also economical. This may also help to increase your understanding of how the WHERE statement, and the SQL GROUP BY, and HAVING clauses work.

147 : The Everytown Research database: Using SAS® analytic procedures to analyze mass shootings
Jayanth Iyengar
Thursday, September 05, 2024, 08:30 AM – 09:00 AM, Location: Regency D
Download White Paper (PDF)
Download Slide Presentation (PDF)

With mass shootings occurring every week, It can accurately be stated that mass shootings in the U.S. have reached the level of an epidemic. Everytown Research and Policy conducts independent methodically rigorous research and supports evidence-based policies to reduce the incidence of gun violence. In 2009, Everytown Research started assembling a Mass Shooting database which records key data on every mass shooting in the U.S. In this paper, I examine and explore the database using SAS® procedures to produce a series of tables, reports, graphics and visualizations. The goal of this project is to generate insights from the SAS analytics that guide the building of effective programs and policies to reduce the epidemic of mass shootings.

148 : Uniform Hashing of Arbitrary Input into Key-Exclusive Segments
Paul Michael Dorfman
Thursday, September 05, 2024, 08:30 AM – 09:30 AM, Location: Regency E
Download White Paper (PDF)
Download Slide Presentation (PDF)

Aggregating or combining large data volumes can challenge computing resources. For example, the process may be hindered by the system limits on utility space or memory and, as a result, either fail or run too long to be useful. It is a natural inclination to try solving the problem by segregating the input records into a number of smaller segments, processing them independently and combining the results. However, in order for such a divide-and-conquer tactic to work, two seemingly contradictory criteria must be met: First, to aggregate or combine the data correctly, no segment can share its key values with the rest; and second, the segments must be more or less equal in size. In this presentation, we show how a hash function can be used to achieve it for arbitrary input with no prior knowledge of the distribution of the key values among its records. Effectively, the method renders any task of aggregating or combining data of any size doable by splitting its input into a large enough number of segments. Such an approach can be used to process the segments sequentially or in-parallel. The trade-off is the need to partially re-read the data. However, it is a rather small price to pay for making a failing or endlessly running task finish on time.

149 : Putting Power into the Hands of the Programmer with SAS Viya Workbench
Joe Madden
Thursday, September 05, 2024, 08:30 AM – 09:30 AM, Location: Golden State
Download Slide Presentation (PDF)

SAS Viya Workbench is available now to empower you with a flexible on-demand compute environment. Regardless of your preferred programming language (SAS, Python and soon R), SAS analytics are embedded within Workbench so you can start building immediately. Whether you have existing SAS 9 code and want to build with the procs you love, or you want to try the newest analytical techniques, Workbench is here to inspire your next innovative breakthrough. This presentation will walk through the first release of Workbench, which is designed for customers who need to keep their data in a private cloud.

150 : Bessler’s Principles of Communication-Effective Data Visualization
LeRoy Bessler
Thursday, September 05, 2024, 10:00 AM – 11:00 AM, Location: Regency D
Download Slide Presentation (PDF)

Let me show you how to best get beyond your graphics software defaults. My list of principles is a L O N G list, but this presentation is a short list. The long list is from 44 years as a data artist working to get the best out of SAS® graphics software. (In the words of Alan Bates, “Just think of it as a hobby that has gotten out of hand.”)

Come to see a subset story, demonstrated with widely applicable data graphics examples that you, too, can use. All, but a dangerous but popular archaism, are rendered with SAS ODS Graphics, The Graphics SuperPower Tool.

If you have SAS software, you inherently have ODS Graphics, at no extra charge. You can use my principles for graphic design and color use with ANY software. (This topic features principles from the book Visual Data Insights Using SAS ODS Graphics: A Guide to Communication-Effective Data Visualization.)

As the world’s longest serving advisor to SAS users on best practices for graphic design and color use for data visualization, I would like to share my ideas, methods, and experiences with you.

151 : Exploration and Revelation for COVID-19 Data: An Atlas and Other Visual Data Insights
LeRoy Bessler
Wednesday, September 04, 2024, 04:00 PM – 04:30 PM, Location: Regency D

A Tableau dashboard posted in the Data Visualization Group at LinkedIn was so underinforming and visually disappointing that I decided to see what could be done better with the superpower tools of ODS Graphics and SAS® software. I applied my (software-independent) principles of communication-effective use of graphics tools and color to real data, not my usual data workspace of SASHELP sample data sets. Let me show you the results for COVID-19 data for 2020 to 2023. All code and input data is available post-conference upon request.

152 : The SAS Supervisor
Don Henderson
Wednesday, September 04, 2024, 04:00 PM – 05:00 PM, Location: Regency A
Download White Paper (PDF)

How SAS processes jobs is the responsibility of the SAS Supervisor and an understanding of its function is important.

While the details of how it works have changed over time, much of the basics of the SAS Supervisor have been reasonably consistent over time.

This paper was originally presented many, many SUGIs ago, specifically, SUGI 83 in New Orleans. It was one of the presentations at the very first Tutorials section. It has been available online as a scanned image thanks to NESUG and is currently available at https://communities.sas.com/t5/SAS-Communities-Library/The-SAS-Supervisor/ta-p/429216

153 : Array Hashing: Simple, Fast, and Efficient
Paul Michael Dorfman
Thursday, September 05, 2024, 04:30 PM – 05:30 PM, Location: Regency C
Download White Paper (PDF)

In SAS programming, using hashing for memory-resident storage and lookup began in 1998 with array implementations of hash algorithms in the SAS language. They vastly outperformed other same-purpose SAS methods available at the time, and so SAS programmers began to include them in their repertoires. However, with the advent of the SAS hash object in 2003, array hashing started falling into obscurity. One, valid, reason is the true power of the hash object. The other is the proliferation of the baseless notion that array-based hash code is too complex to comprehend and maintain. This is quite unfortunate: Under many scenarios, array-based hash search is simpler, much faster and more efficient than the hash object. This paper presents the simplest (and yet most practical) hash search algorithm and its array-based SAS implementation as plain modifications of the sequential search. We will also see examples of how array-based hashing can be used to perform basic table operations (such as Search, Retrieve, Insert, Update, etc.) and how it performs vis-à-vis the hash object in terms of speed and efficiency.

155 : Leading with Impact: Building Stronger Programming Teams Through Effective Leadership and Collaboration
Menaga Guruswamy Ponnupandy
Thursday, September 05, 2024, 10:30 AM – 11:00 AM, Location: Regency C
Download White Paper (PDF)

In the fast-paced environment of Contract Research Organizations (CROs), programming teams are essential for crafting innovative solutions and ensuring operational efficiency. However, these teams often face challenges that hinder productivity and job satisfaction, including issues with resource allocation, overtime management, and conflict resolution. This white paper investigates these challenges and proposes practical solutions to enhance a cooperative and dynamic work environment. By focusing on enhancing professional development, adopting inclusive leadership practices, managing workloads effectively, and mitigating micromanagement, the paper aims to help CROs build a culture of collaboration and empowerment. The paper is structured around ten detailed case studies, each providing specific scenarios and solutions designed to enhance team performance, increase employee retention, and drive sustainable growth and innovation. These improvements are intended to guarantee the company’s long-term success.

157 : Will statistical programmers be replaced by Artificial Intelligence?
Junying Zhang
Friday, September 06, 2024, 10:00 AM – 10:30 AM, Location: Regency D
Download Slide Presentation (PDF)

Will statistical programmers be replaced by Artificial Intelligence?

Last year, I shared my working experience as a statistical programmer (SP) in social media, and I was challenged by someone using the ChatGPT. And he declared that in the future ChatGPT will create the TFLs from statistical analysis plan and CRF rawdata and incorporate those into the clinical study report directly. And statistical programmers and medical writers will be replaced by artificial intelligence (AI) soon.

As a SP in the pharmaceutical industry for 10 years, I think about my daily work and the below questions. Whether SP’s all tasks could be replaced by Artificial Intelligence? Whether all SP’s work will be standardized?

In fact, there are a lot of manual review work for statistical programmers, and also many collaborations and communications with many other functions within study team. And it does need a lot of industry working experience to be a good statistical programmer.

There are some essential SP tasks that AI could not replace an experienced programmer. Examples include study protocol and statistical analysis plan review, EDC set up review, CDISC standards adherence, data issue resolution, and communicating with cross-functional teams and external vendors. Statistical programmers play a vital role in ensuring the accuracy and integrity of data transformation processes, such as mapping CRF and vendor data to Study Data Tabulation Model (SDTM), and creating Analysis Data Model (ADaM) datasets for analysis according to CDISC standards.

From the above, a good statistical programmer does much more than just running the SDTM, ADaM, and TFLs outputs. There are a lot of manual review work behind the datasets and TFLs. While AI and standardization can automate certain tasks and streamline processes, it cannot replace the critical thinking, problem-solving, and communication skills that statistical programmers bring to the table.

159 : Estimating causal effects of health interventions using instrumental variables
Mehrnaz Siavoshi
Thursday, September 05, 2024, 09:00 AM – 09:30 AM, Location: Regency D
Download White Paper (PDF)
Download Slide Presentation (PDF)

Consistent estimation of causal effects is pivotal in health research to inform effective policy-making and treatment strategies. This paper presents a brief exploration of instrumental variables (IVs) as a methodological approach for estimating causal effects when randomization is impractical or not feasible, such as when utilizing observational data. In this technique, traditionally employed in the fields of epidemiology and econometrics, a third endogenous variable is incorporated in model development to understand the causal relationship between two endogenous variables, while allowing for controls and other traditional modeling techniques. The overarching goal of this technique is to use a variable correlated with the treatment, but not with the outcome except through the treatment variable, to isolate exogenous variation in the treatment and therefore allowing for greater precision in causal estimation. This paper will give the rationale behind IVs and methodologies for selecting proper variables will be examined. SAS procedures for IV analysis will be provided, including PROC SYSLIN, with a focus on model specification, instrument selection, and interpretation of results. Recommendations for instrument selection, special considerations, and challenges in using IVs will also be included.

160 : iCSR: A Wormhole to Interactive Data Exploration Universe
Sudhir Kedare, Steve Wade
Thursday, September 05, 2024, 03:00 PM – 03:30 PM, Location: Regency C
Download White Paper (PDF)
Download Slide Presentation (PDF)

In the past decade, CDISC has brought standardization to clinical data, significantly benefiting clinical
development by allowing regulatory reviewers to review submitted data with efficiency and precision.
Despite these advances, bottlenecks such as static TFL report submission hinder the overall efficiency of the review cycle.

At Jazz, we are continuously seeking innovative ways to explore data to establish a foundation for digital transformation. Namely, we have developed the Interactive Clinical Study Reports (iCSR) application, utilizing CDISC-standardized data and RShiny capabilities to enable interactive data exploration.

This paper demonstrates how iCSR facilitates real-time dynamic data exploration and significantly
improves the data review process by providing fast insights and enabling informed decision-making.
iCSR comes with an intuitive UI and minimal onboarding. For pharmaceutical organizations, iCSR is a
promising solution to optimize data review and lead efforts to get effective therapies in patients’ hands as early as possible.

161 : inspectoR: QC in R? No Problem!
Steve Wade, Sudhir Kedare
Friday, September 06, 2024, 10:30 AM – 11:00 AM, Location: Regency D
Download White Paper (PDF)
Download Slide Presentation (PDF)

More organizations are starting to embrace open-source technologies to perform tasks traditionally completed in SAS®. One such activity is to QC datasets, tables, and figures in the process of producing TLF’s. Independent programming is done for many of those TLF’s, comparing the results from both programmers. Jazz has developed the inspectoR package, an alternative to the SAS® COMPARE procedure, to allow QC performed in R to be compared back to datasets generated using SAS®.

In this paper, we demonstrate how inspectoR will compare these datasets and produce a report showing the findings. The report produced by inspectoR is much like PROC COMPARE output but is produced using HTML in a more readable format.

inspectoR has proven to be a valuable tool in helping to transition QC tasks to R and maintain the level of quality expected from SAS® systems.

162 : Interactive Data Analysis and Exploration with composR: See the Forest AND the Trees
Steve Wade, Sudhir Kedare
Friday, September 06, 2024, 10:00 AM – 10:30 AM, Location: Regency C
Download White Paper (PDF)
Download Slide Presentation (PDF)

Deciding on Post-hoc analyses can be a time-consuming process and it’s critical an analysis has been vetted prior to production and release. Jazz is pioneering creative techniques in data discovery using open-source technologies. Accordingly, composR is an interactive R-Shiny application that provides a user-friendly interface for rapid data analysis and exploration.

This paper demonstrates how composR allows users to filter data, create new variables, and perform analyses, summaries, and visualizations. It also enables further drill-down into summaries and charts for more in-depth analysis. Although not required, composR supports CDISC standard data from a variety of file types including SAS®, XPT, CSV, Excel, and Rda.

ComposR reduces time, effort and cost over traditional TLF processes by providing insights into analyses prior to formal TLF production. This ensures only necessary analyses are produced, eliminating much of the back-and-forth between departments.

166 : Statistical Programmers Role for SPM Analysis
Marckenley Mercie, Ce (Colin) Zhou
Thursday, September 05, 2024, 03:30 PM – 04:00 PM, Location: Regency E
Download White Paper (PDF)

In the context of multiple myeloma (MM) clinical trials, analyzing secondary primary malignancies (SPM) is crucial. Regulatory obligations require thorough investigation and reporting of SPM occurrences. However, a lack of consensus on executing this analysis leads to industry variations. This paper outlines a concise process for creating analysis datasets and offers recommendations for standardizing SPM reports to support this critical analysis. It also emphasizes the essential role of statistical programmers in conducting these analyses effectively.

169 : Customizing SAS Studio in Viya: The Next Step
Danny R Modlin
Thursday, September 05, 2024, 11:00 AM – 11:30 AM, Location: Golden State
Download Slide Presentation (PDF)

SAS has always been there for users to assist with organizing and developing code. This demonstration discusses the next tool for such assistance in statistical analysis, the custom step. Custom steps enable you to create a user interface for a specific task or analysis. Add your custom steps to a SAS Studio flow to streamline and automate your statistical analysis! In this demo, you will learn to create your own statistical custom step in the SAS Studio Designer.

171 : Identifying a Skilled Nursing Facility Associated COVID-19 Case in Surveillance Data
Kathryn Doh, Maggie G Tufts
Thursday, September 05, 2024, 09:30 AM – 10:00 AM, Location: Regency D
Download White Paper (PDF)
Download Slide Presentation (PDF)

Skilled Nursing Facilities (SNFs) posed a unique public health concern during the COVID-19 pandemic, having both an elevated risk of transmission, as a setting where individuals lived and worked with limited ability to isolate, and an elevated risk for severe disease and death, as a setting serving medically vulnerable individuals. To provide as accurate as possible SNF COVID-19 case and death counts, we developed a series of macros to identify cases likely occurring in SNF residents or staff using a combination of free-text search, cross-walked selection fields, and address matching. Results from this effort were used to validate facility reported data, get demographic data, and identify cases not linked to an outbreak that likely should be.

172 : How to Modify SAS 9 Programs to Run in SAS Viya
Danny R Modlin
Thursday, September 05, 2024, 03:30 PM – 04:00 PM, Location: Regency C
Download Slide Presentation (PDF)

How can existing SAS 9 programs can be modified to execute in SAS Viya. Code can either run as is on the SAS Compute Server, or it can be modernized to process data in memory and in parallel on the SAS Cloud Analytic Services (CAS) server. This presentation is perfect for programmers who are new to SAS Viya and want to continue performing their statistical analyses there. We will address questions that are typically asked. 1. Will existing SAS 9 code work in Viya? 2. How must my programs change to take advantage of the new features in Viya?

174 : X-Ray Image Classification with Neural Networks
Riley Rutan, Eric C Braga, Xinyu Du
Friday, September 06, 2024, 10:00 AM – 10:30 AM, Location: Golden State
Download White Paper (PDF)

Machine learning has the potential to revolutionize the way that medical professionals
review medical results to make diagnoses. A model that can classify large batches of medical
imaging results to identify medical conditions and diseases can greatly improve patient outcomes
as well as the efficiency and accuracy of radiologists and healthcare providers. Here, we created
a neural network model that utilized deep learning techniques to analyze and classify medical
images to identify a variety of diseases. A tool that can leverage machine learning to help
medical professionals identify diseases more quickly and reliably will be an invaluable resource.
We have focused our model on categorizing chest X-ray images to identify respiratory
diseases such as COVID-19, viral and bacterial pneumonia, and tuberculosis. We have utilized
neural networks that can identify and map the complex and subversive patterns that exist in the
data stored in medical images. We leveraged the Keras interface of TensorFlow, a machine
learning and AI library, to build models that can quickly and accurately categorize these images
at a success rate that meets industry standards. While implementing this technology raises
ethical and privacy concerns, the use of neural networks undoubtedly drives progress in future
healthcare.

175 : Flight Delay Analysis and Prediction
Eric C Braga, Riley Rutan, Miguel Angel Bravo Martinez del Valle, Xinyu Du, Fernanda Carrillo
Friday, September 06, 2024, 11:00 AM – 11:30 AM, Location: Regency D
Download White Paper (PDF)

Like most global industries, the COVID-19 pandemic caused fundamental shifts in the airline
industry. As a result, the consumer is often left with the inconvenience of delayed and canceled flights
for various reasons. In this project, we explored a large dataset of canceled and delayed flights from
2019-2023 from the US from the Department of Transportation along with weather data from NOAA,
leveraged machine learning models to identify trends in delayed and canceled flights, and created a tool
that consumers can use to predict the probability that their upcoming flight will be delayed or canceled.
After using several machine learning models, we utilized an Extreme Gradient Boosting (XGBoost) that
achieved an accuracy of 0.68 and a weighted average f1-score of 0.70.

179 : Training Domain Specific AI for Clinical Data: Developing an AI Agent to Transform EDC data to CDISC SDTM and then to ADaM Standards
Sy J Truong
Thursday, September 05, 2024, 04:00 PM – 04:30 PM, Location: Regency E
Download White Paper (PDF)

The transformation of clinical trials data from Electronic Data Capture (EDC) systems into CDISC-compliant formats, such as the Study Data Tabulation Model (SDTM) and the Analysis Data Model (ADaM), is resource-intensive, requiring significant manual effort and expertise. Automating this transformation with a domain-specific AI model provides efficiencies and reduce errors. This paper explores the development and training of such an AI model, utilizing SAS as the primary scripting language, complemented by Python for model training and integration with open-source language models (LAMA) and OpenAI’s latest models.

The proposed approach is to fine-tune a pre-trained model with domain specific data by incorporating domain knowledge and rules through iterative training and refinement. The system uses SAS for robust data handling and integration within the clinical data workflow, while Python facilitates AI model training. Our methodology leverages LAMA and OpenAI’s models to enhance the AI’s understanding and generation of transformation rules to gain efficiencies.

The training process begins by leveraging existing SDTM mapping knowledge and scripts to convert EDC data to SDTM format. Domain-specific rules are applied to transform the data into SDTM format, followed by fine-tuning through semantic parsing of SAS programs. These rules, encoded in SAS scripts, serve as training data for the AI model, which learns to predict appropriate transformations based on input data characteristics.

To transition from SDTM to ADaM, the AI model is further trained on curated training data to understand analysis-specific requirements and apply necessary transformations. The model is fine-tuned using supervised sequence-to-sequence modeling and semantic parsing to comprehend statistical analysis.
The combination of SAS scripting and Python-assisted AI training presents a powerful approach to automating clinical data transformation, significantly improving efficiency and accuracy. This work lays the foundation for future advancements in AI-driven automation for mapping clinical trial data to domain-specific standards.

182 : Analyzing Goalkeeper Impact in Premier League Football
Alec Troy Christiansen, Benjamin T Houser
Friday, September 06, 2024, 11:30 AM – 12:00 PM, Location: Regency D
Download White Paper (PDF)

The sport of football (US “soccer”) contains untapped potential in the world of data analytics. The sport’s main objective is scoring goals, and one position plays a pivotal role in whether goals are scored or not. The goalkeeper role is widely considered one of the most difficult positions to play in football. Knowing the key aspects of a goalkeeper’s performance is vital to understanding the success of a club. This study seeks to quantify the importance of a goalkeeper’s ability to perform aspects of his role as it holistically impacts the team’s performance within the context of the English Premier League. The analysis will utilize data from Sports Reference’s Premier League raw data on goalkeeper statistics for the five most recently completed seasons (2018-2023) as well as wins, losses, and draws recorded by clubs. We conducted stepwise regression techniques to identify key variables and construct a model based on the results with the lowest AIC values. Our null hypothesis motions that there will be no goalkeeping statistics that are significant for predicting club wins in a season. However, we found that Save Percentage and Opponent Crosses were highly significant for club win prediction in the Premier League from 2018-2023. The findings of this analysis will inform clubs looking to optimally develop their coaching and recruiting to prioritize winning games in the Premier League and in all of football.

183 : A SAS® Macro to Calculate Therapeutic Intensity Score
Rachelle Juan
Friday, September 06, 2024, 09:30 AM – 10:00 AM, Location: Regency E
Download White Paper (PDF)

Hypertension is the leading cause of cardiovascular disease and premature death globally. Therapeutic inertia, or the failure to initiate or intensify therapy when treatment goals are not met, has been identified as a key barrier to effective hypertension control. Evaluation of therapeutic intensification by calculating the therapeutic intensity score (TIS) can help to raise awareness, monitor, and properly manage uncontrolled blood pressure in a timely manner. Despite the importance of TIS in hypertension treatment, its use is limited due to the lack of standardized programs or tools for TIS calculation. This paper presents a systematic approach and a simple SAS® program demonstrating the data structure and programming process for calculating the TIS and defining therapeutic intensification of hypertension medications. To demonstrate our methods, we used pharmacy data from the electronic health records of Kaiser Permanente Southern California. We believe that the application of this standardized procedure and SAS program will help increase the utilization of TIS, thereby addressing therapeutic inertia and improving hypertension control.

185 : SAS Viya vs SAS 9: Architecture and Performance For SAS Users / Teaming With Your SAS Architect and Admin For A Fast SAS Viya
Michael Shealy
Friday, September 06, 2024, 09:30 AM – 10:00 AM, Location: Regency C
Download Slide Presentation (PDF)

SAS Viya is the flagship analytics platform of the SAS Institute. Optimized for deployment on cloud, the platform brings many changes from the previous SAS 9 architecture. When it comes to performance optimization, however, many of the principles remain the same across both offerings. This presentation provides an overview of the logical and physical architectures of SAS Viya and SAS 9, shows how many of the performance drivers between the two platforms remain the same, and where the Viya platform may bring differences. It wraps up with how SAS Users can team with their SAS Architects and Administrators to help optimize the performance of their Viya platform.

186 : A Framework for Clinical Trial Data Synthesis
James Joseph
Thursday, September 05, 2024, 02:00 PM – 03:00 PM, Location: Golden State

This paper introduces a framework for synthesizing clinical trial data from real study protocols, showcased by our Synthetic Data Definition Table (sDDT), adapted from traditional Data Definition Tables (DDT).

We illustrate how the sDDT leverages industry standards, defines precise expectations for study-specific data synthesis, and ensures reproducibility.

We demonstrate the data synthesis in a training scenario where statistical programmers manage changes to source data extracts, thereby deepening their understanding of clinical trial concepts and gaining practical experience.

We discuss the potential for this framework to improve and expand to applications such as assessing data quality and developing external control arms.

187 : Simulation in SAS for Estimating Power
Michael Williams
Wednesday, September 04, 2024, 02:30 PM – 03:00 PM, Location: Golden State
Download Slide Presentation (PDF)

Simulation is a great way to estimate key quantities in statistics. Given limited time and resources, exact calculations may not be feasible and/or understandable. This presentation will focus on the techniques of Monte Carlo simulations for power of a hypothesis test. The running example will be the estimation of power of the single sample t-test via simulation, with a comparison to the exact calculation from PROC POWER.

188 : Simulating Optimal Sample Sizes for Joyful Canine Jaws Using SAS
Chary Akmyradov, Lida Gharibvand
Thursday, September 05, 2024, 11:00 AM – 11:30 AM, Location: Regency D
Download White Paper (PDF)
Download Slide Presentation (PDF)

In the realm of veterinary clinical research, ensuring animal welfare while achieving statistically significant results is paramount. This study presents an advanced approach to sample size calculation for a clinical study involving canine subjects, with a focus on dental health. The unique design of this study employs dogs as both cases and controls by longitudinally comparing treated and untreated teeth within the same animal, thus minimizing the number of subjects required and reducing animal suffering.

A pivotal aspect of this research is the optimization of the number of dogs and the number of teeth extracted per dog. The goal is to minimize both, ensuring minimal discomfort to participating animals. The teeth growth in dogs is monitored at three distinct time points to assess the development of treated versus untreated teeth.

To achieve a robust and reliable study design, a simulation-based approach was adopted. This involved simulating canine teeth growth trajectories based on pilot studies and existing literature using a Data Step. Power analysis was conducted using the simulated data, utilizing the PROC GLIMMIX and PROC FREQ procedures. Additionally, the entire simulation process was streamlined and automated using a custom SAS macro. Lastly, the results are visualized as a heat map using PROC SGPANEL.

This paper highlights the delicate balance between ethical considerations and the need for scientific rigor in veterinary research. The methodology outlined here serves as a blueprint for future studies requiring minimal animal subjects while ensuring reliable and ethically sound outcomes.

189 : The Business Value of Diversity Combined with Data Science
Stephen Sloan
Thursday, September 05, 2024, 04:00 PM – 04:30 PM, Location: Regency C
Download White Paper (PDF)

Having strong diversity programs, being sensitive to diversity issues, and being able to apply rigorous data science techniques can provide significant value to an organization.
Many business problems and challenges begin with the statement of a problem, and the problem sometimes seems to relate to issues with diversity. Organizations are very aware of the importance of diversity and problems sometimes manifest themselves as diversity issues, even when they have other causes. At other times there are issues where awareness of the value of diversity can help an organization achieve its goals. In addition, there are areas where diversity can have a direct impact on the organization in terms of compliance and scientific value.
In this paper I cite examples where diversity has proven to have business value for an organization.

190 : Regression Analysis Made Easy Using SAS® Studio
Zheyuan Yu
Wednesday, September 04, 2024, 04:00 PM – 05:00 PM, Location: Regency E
Download White Paper (PDF)
Download Slide Presentation (PDF)

SAS® OnDemand for Academics (ODA) provides students, faculty, and SAS learners with free access to SAS software and the SAS® Studio user interface using a web browser. SAS Studio provides a comprehensive and customizable integrated development environment (IDE) for all SAS users. To showcase SAS Studio’s many features, numerous techniques will be introduced to access, clean, transform, analyze, and visualize data using the point-and-click features found in SAS Studio’s Navigation Pane’s Tasks and Utilities. Plus, we’ll demonstrate the generated SAS code that is automatically produced from the point-and-click techniques. To obtain a high-level understanding of the datasets being used, we’ll demonstrate tasks associated with exploratory data analysis (EDA) to identify missing values, explore outliers, and evaluate trends in the data. Two types of regression will be demonstrated – simple linear regression where one independent variable is used to explain or predict the outcome of the dependent variable and multiple linear regression where two or more independent variables to explain or predict the outcome of the dependent variable to assist with decision-making activities. Key takeaways will be provided to assist in learning regression analysis techniques using effective examples.

191 : Developing a SAS Macro for ISNI: Assessing Missing Data Sensitivity Beyond MAR Assumptions
Bocheng Jing, L. Grisell Diaz-Ramirez, John Boscardin
Thursday, September 05, 2024, 11:30 AM – 12:00 PM, Location: Golden State
Download White Paper (PDF)

Missing data is a common challenge in research, often addressed by statistical methods assuming that the data are Missing At Random (MAR). While this assumption simplifies analysis, it is unverifiable and can introduce bias if incorrect. To address this, sensitivity analyses assess the robustness of conclusions to deviations from MAR. The Index of Local Sensitivity to Nonignorability (ISNI), introduced by Troxel et al. (2004) for cross-sectional data and extended by Xie et al. for longitudinal data, provides a straightforward method for this purpose without requiring complex non-MAR models.

Currently, the ISNI method is available in the R package “isni” for various regression models, featuring user-friendly syntax. However, a corresponding SAS procedure has yet to be developed. To fill this gap, we present a SAS macro, %isni, for calculating the ISNI index. This macro enables sensitivity analysis for multiple linear regression, logistic regression, Poisson regression, and negative binomial regression in the presence of missing outcomes.

By introducing the %isni macro, we bring ISNI methodology to the SAS community, facilitating robust sensitivity analysis for researchers using SAS. This development aims to encourage the incorporation of ISNI into standard SAS procedures in the future.

193 : Comparative Examples of Geocoding and Mapping Techniques in SAS® and R Using Centers for Disease Control Survey Data
Joshua J Cook, Louise S Hadden, Swann Arp Adams
Thursday, September 05, 2024, 03:30 PM – 04:30 PM, Location: Regency D
Download White Paper (PDF)

The Centers for Disease Control (CDC) maintains the Behavioral Risk Factor Surveillance System (BRFSS), which is a rich source of survey data for public health and epidemiological research. BRFSS collects data in all 50 states as well as the District of Columbia and three U.S. territories, and includes data on health-related risk behaviors, chronic health conditions, and use of preventive services. Geospatial analysis is a crucial tool in public health and epidemiology, offering insights into geographic patterns and trends. This presentation aims to provide a comprehensive comparison of geocoding and mapping techniques using SAS® and R, utilizing data from the CDC BRFSS. By demonstrating how to geocode and visually map public health data, we seek to empower researchers and practitioners with the skills necessary to conduct spatial analyses. In this session, we will illustrate the process of geocoding and mapping BRFSS data using SAS® PROC GEOCODE and PROC SGMAP procedures, alongside equivalent methods in R, using packages such as tidygeocoder, ggplot, and tmap. Attendees will gain insights into the strengths and limitations of each approach, understanding how to choose the appropriate tools for their specific needs. Key topics will include an overview of the CDC BRFSS dataset and its relevance to public health, step-by-step guidance on geocoding and generating spatial coordinates, techniques for creating informative and aesthetically pleasing maps, with direct comparisons between SAS® and R in terms of functionality, ease of use, and output quality. By the end of this lecture, participants will be equipped with practical knowledge to apply geocoding and mapping techniques in their work, enhancing their ability to analyze spatial data and make data-driven decisions in public health.

194 : Qualitative and Quantitative Data Analysis and Visualization Using Python
Leon Rod Davoody, Lida Gharibvand
Thursday, September 05, 2024, 09:30 AM – 10:00 AM, Location: Regency C
Download White Paper (PDF)
Download Slide Presentation (PDF)

Abstract: This paper explores the capabilities of Python in both qualitative and quantitative data analysis and visualization. From data processing to visualization techniques, Python offers a versatile toolkit for researchers and analysts working with diverse datasets. The paper delves into practical examples, showcasing the application of Python libraries for qualitative and quantitative analysis.
Introduction: Python has become a prominent language for data analysis and visualization, offering a comprehensive ecosystem of libraries. This paper provides an overview of Python’s capabilities in handling both qualitative and quantitative data, demonstrating its utility in research and decision-making processes.

Method: Python has been used to visualize qualitative and quantitative data analysis. – Using bar charts and pie chart to visually represent qualitative insights. Also, using histogram and box plot to visually represent qualitative insights.

Conclusion: Python’s rich ecosystem makes it a powerful tool for both qualitative and quantitative data analysis and visualization.

196 : From Keys to Credit: A Deep Dive into Homeownership’s Influence on Loan Quality
Yi-Che Chen, Debanik Chakraborty
Thursday, September 05, 2024, 11:30 AM – 12:00 PM, Location: Regency D
Download White Paper (PDF)
Download Slide Presentation (PDF)

In banking, analysts strive to identify traits predicting loan delinquency. Homeownership status, reflecting stability and financial behavior, poses a crucial question: “Does homeownership significantly influence loan delinquency?” The objectives of this study were twofold: building a predictive model leveraging borrower attributes to assess loan quality and interpreting homeownership’s impact. To achieve this, we aimed to combine real-world public data sources with the loan dataset collected from Kaggle, which contains 27 attributes 396,030 individual. We reviewed the data and performed data manipulation to create new variables. To address imbalanced data, we applied SMOTE sampling technique to train our models and tested on the original data to check the practical implication of MLR models. Six classification prediction models (Logistic Regression, Decision Trees, SVM, Neural Networks, Random Forests, Gradient Boosting) were implemented. These models were evaluated using Accuracy, F-1 Score, ROC-AUC score, KS (Youden) and Misclassification Rate. The final model aims to inform credit assessment, providing insights into homeownership’s role in loan quality. Logistic Regression proves to be the most accurate model with the highest ROC-AUC (71.46%), F-1 Score (43.43%) and one of the lowest False Negative Rate (31.64%). We found total interest rate, term of the loan, debt to income ratio, income verification status and revolving credit utilization rate to be the most important features. Our model suggests that the chance of delinquency is more for the applicants without a mortgage house, though it has no impact on the applicant having a mortgage house.

199 : Advanced DATA Step Look-Back and Look-Ahead Techniques
Josh Horstman
Thursday, September 05, 2024, 04:30 PM – 05:30 PM, Location: Regency E
Download White Paper (PDF)

The DATA step is a powerful and versatile tool for manipulating SAS data sets. Because it is built around the concept of reading one record at a time into the Program Data Vector, DATA step programming logic generally has access to only one record at a time. However, there are many situations in which it is advantageous to be able to look back at values from prior records or to look ahead at values from later records. This presentation introduces several advanced DATA step programming techniques that allow the programmer to work with values from multiple records concurrently. Approaches discussed include the double SET statement, the self-merge, the LAG function, and several others. A number of examples are used to explain each method and demonstrate its use.

202 : Count, Lag, Retain! Let’s go through it one more time!
Steve Black
Wednesday, September 04, 2024, 04:30 PM – 05:00 PM, Location: Regency D
Download White Paper (PDF)
Download Slide Presentation (PDF)

In SAS there are few things that can turn a good programmer into a pale, sleep deprived, shadow seeking individual like a RETAIN statement that is not working right. In this paper I hope to provide a deeper understanding how to better count on and count with the RETAIN statement in SAS. I will illustrate a number of ways to count using the properties of the program data vector and provide some real life examples of how I have found using the RETAIN statement super helpful. Coupled with these two options I’ll also touch on the LAG function. In the end I hope to bring out of the shadows and into the light the RETAIN, count and LAG functions.

203 : It’s All about the Base—Procedures
Jane Eslinger
Wednesday, September 04, 2024, 03:30 PM – 04:00 PM, Location: Regency E
Download White Paper (PDF)

As a Base SAS® programmer, you spend your day manipulating data and creating reports. You know there is a procedure that can give you what you want. As a matter of fact,there is probably more than one procedure to accomplish the task. Which one should you use? How do you remember which procedure is best for which task?

This paper is all about the Base procedures. It explores the strengths of the commonly used, nongraphing procedures. It discusses the challenges of using each procedure and compares it to other procedures that accomplish similar tasks. The first section of the paper looks at utility procedures that gather and structure data: APPEND, COMPARE, CONTENTS, DATASETS, FORMAT, SORT, SQL, and TRANSPOSE. The next section discusses the Base SAS procedures that work with statistics: FREQ, MEANS/SUMMARY, and UNIVARIATE. The final section provides information about reporting procedures: PRINT, REPORT, and TABULATE.

205 : Taking the Mystery Out of and Debugging PROC HTTP
Kim Wilson
Friday, September 06, 2024, 08:30 AM – 09:30 AM, Location: Regency C
Download White Paper (PDF)

Several great papers have been written about how to get started with PROC HTTP, which includes accessing Microsoft 365 applications, modifying various options for desired results, and more. As a SAS Technical Support Engineer, I often assist SAS customers who are not receiving the expected resource, or they are seeing a return code that is not a 200 OK. This paper describes common errors that you might encounter regarding certificates, authentication, and general errors, as well as overall debugging techniques and suggestions. This paper also helps you gather pertinent information that SAS Technical Support will need when helping to solve the problems occurring with or around PROC HTTP.

208 : Customized LOG/LST name with the SAS® Config File
Zeke Torres
Thursday, September 05, 2024, 10:30 AM – 11:00 AM, Location: Regency E
Download White Paper (PDF)

The SAS® config file is powerful and useful. In this example, we show how to customize the names of the Log/Lst files from the code we run, with a useful new name as the result.
The new log file might look like this:
20180904_hhmm_userID_name-of-code-that-was-run.log.
The benefit is that when one user or a team of users builds the code, everyone can see its progress. In this way, collaborating is simplified and results are easier for colleagues to share and consume. This is a mock or proposed file name. This macro is intended to inspire you to see what additional SAS System values to add to your style.

The typical problem encountered is we can run our SAS code and get useful output: LOG/LST – but after we run it again – that original output is gone. Its overwritten.
Or if one of our co-workers or team members runs the same code – again – now our LOG/LST become harder to pin down – who – ran that code. This is a solution meant to utilize the SAS Configuration file and a set of SAS Macros to improve the name of the LOG/LST output. There are already ways to customize those – with SAS options. This is more of a recommended methodology for you to consider. My reason for employing this technique is because often I am searching for information in many LOG/LST. Either to debug, help someone debug (Myself, Team Member, Client) or to track progress on a “Project/Build” of work that is critical to gauge where and how things are. Especially when one or more work on something like this.

210 : Multiple Imputation in Interim Analyses
Bill Coar
Wednesday, September 04, 2024, 04:30 PM – 05:00 PM, Location: Golden State
Download Slide Presentation (PDF)

In adaptive designs, formal interim analyses may be performed in order to alter the study based on an interim review of data. This includes calculation of conditional power, futility analysis, and early efficacy analysis. In most cases, the statistical approaches using the interim data will be consistent with those planned on the final data. Sometimes the approaches are straightforward while others such as the use of mixed models repeated measure and multiple imputation are more complicated. Missing data can be imputed using multiple imputation (MI) techniques. MI techniques are typically pre-specified but missing data patterns will be a function of the data. Additionally, different programming approaches (such as the use of single or multiple calls to Proc MI) will result in different imputation values due to the random nature of the imputations. The imputation dataset will contain replicates of the original dataset but with the missing values populated. Each replicate is analyzed, and final results averaged together for a single final estimate for decision making using Proc MIANALYZE.
This presentation will introduce MI and associated assumptions, provide an example, and discuss various challenges that can occur when analysis is performed in interim data.

214 : The Dummy’s Guide to Getting Started with SAS
Chris Hemedinger
Thursday, September 05, 2024, 02:00 PM – 03:00 PM, Location: Regency D
Download Slide Presentation (PDF)

This lecture is designed for beginners who want to learn the basics of SAS software and programming. By the end of this presentation, you will be able to: Understand what SAS is and how it works; Navigate the SAS user interface(s) and access data sources; Write and run SAS programs using basic syntax and logic; Perform common data manipulation and analysis tasks; Generate and customize reports and graphs; and Use SAS online resources and communities to enhance your skills.

215 : Methods and Tools for Publishing SAS Papers, Books, and Other Documentation
Jim Blum, Jonathan W Duggins
Wednesday, September 04, 2024, 05:00 PM – 06:00 PM, Location: Regency A
Download White Paper (PDF)

For programmers, publishing papers or other documents may be a frustrating task. We show some methods of applying programming in LaTeX and SAS to create styles and structures that can be easily reused and revised, with the added benefit of actually looking good. We build styles for various documents elements including boxes/titles for code, output and logs (and other items). LaTeX tools for numbering sections, programs, output, and other elements inside the document are demonstrated along with making cross-references to such items which automatically update when items are moved or removed. An introduction to the ODS TAGSETS.COLORLATEX destination and the resulting sastable object is also provided. Further, automation methods for including SAS code and output are provided with SAS programs for converting .SAS files to .TEX files. As the ODS TAGSETS.COLORLATEX command is pre-production and has limited features, we provide and demonstrate SAS code to modify sizing for tabular output objects based on corresponding RTF sizing. We combine SAS code and output building strategies with LaTeX tools for automatic insertion of the converted SAS code and SAS output into the LaTeX document. All LaTeX implementations use the Overleaf platform, but other LaTeX environments work as well.

220 : Introduction to Conditional Power by Example
Bill Coar
Thursday, September 05, 2024, 09:30 AM – 10:00 AM, Location: Golden State
Download Slide Presentation (PDF)

Clinical trials are designed to have a high probability of detecting a pre-defined treatment effect which is assumed to be true. If a new treatment is truly effective, we desire a high probability of detecting a statistically significant treatment effect at the end of the study. This is referred to as the power of a study. A sample size is calculated to ensure the desired power to detect a treatment effect.
Once a trial is under way, the probability of detecting a treatment effect can be recalculated given the observed data thus far. We can estimate this power conditioned on the observed data once a pre-defined fraction of patients complete treatment and typically assume the remainder of the data will follow the observed trend. This estimate of conditional power (CP) can be extremely informative for decision making when using adaptive designs in clinical trials. Such trials adapt partway through using pre-defined rules. In some cases, these rules are based on CP. We introduce the concept of CP by way of example of testing for a difference between two sample means.

Panel Discussions

106 : Benefits, Challenges, and Opportunities with SAS® and Open-Source Software (OSS) Integration
Kirk Paul Lafler, Ryan Paul Lafler, Joshua J Cook, Stephen Sloan, Anna T. K. Wade
Thursday, September 05, 2024, 04:30 PM – 05:30 PM, Location: Regency D
Download White Paper (PDF)
Download Slide Presentation (PDF)

With all the power and functionality that SAS® has to offer, it now provides users with exciting integration paths to the world of open-source software (OSS). This presentation will create a dialog to engage audience members about strategies, techniques, and integration pathways for integrating the world of SAS software with open-source initiatives and technologies. Technology in the 21st century is experiencing a paradigm shift where organizations around the globe are demanding that all software live, play, and thrive in the same sand box together. We will also discuss the challenges facing the user community as they grapple with methods and approaches to best integrate SAS with open-source software, compatibility and vulnerability issues, security limitations, intellectual property issues, warranty issues, and inconsistent developer practices. Plan to join us for an exciting presentation about the benefits, challenges, and opportunities confronting us all with the integration of SAS software with the incredible opportunities coming from the open-source community, including cloud computing architecture, open standards, and the collaborative nature of community.

197 : What comes after the semicolon? How I learned to stop worrying and love the meeting
Isaiah Lankham, Matthew T Slaughter
Thursday, September 05, 2024, 11:00 AM – 12:00 PM, Location: Regency C
Download Slide Presentation (PDF)

In this panel discussion, we’ll convene a group of experts to help us think through questions like these:

(1) What would your ideal meeting look like?

(2) Is it possible to have too few meetings?

(3) How do you balance meeting overload versus email overload?

(4) Can programming skills help us run more effective meetings?

(5) Does how you think about meetings change in a management role?

(6) What are the challenges for technical employees in management roles?

SAS-quatschen Facilitated Discussions

154 : Get a GPS to Navigate Your Skills to Find Career Purpose
Lisa A Mendez, Charu Shankar
Thursday, September 05, 2024, 11:30 – 12:00, Location: Resource Central
Download White Paper (PDF)
Download Slide Presentation (PDF)

The global pandemic saw employees take charge of their careers and focus on what inspired them in
never seen ways. Finding purpose can seem overwhelming and a daunting task in these changing times. To address this issue, Charu and Lisa have will guide knowledge workers to find purpose and fulfillment in their work while bringing their best strengths forward. Using their combined skills and expertise in psychometrics, SAS, and testing and measurement, they will share self-assessment tools to understand one’s skill set and then apply it to job fulfillment. Each participant will be provided with a self-assessment tool that will be discussed during the presentation.

This is an interactive paper presentation. We invite the attendees to reflect and answer questions while going through the process of navigating to find a career and life purpose.

195 : The Current State of Teaching Biostatistics in Academia
Lida Gharibvand
Wednesday, September 04, 2024, 04:30 PM – 05:00 PM, Location: Resource Central
Download Slide Presentation (PDF)

The traditional approaches to teaching and learning biostatistics have gone through evolutionary stages. This is mainly due to the advancement of computing and modern approaches of data analytics, including but not limited to the methodological impacts statistical machine learning. However, the growth and impact of biostatistics programs, both at the undergraduate and graduate levels, are heavily reliant upon building strong and relevant foundations. It is from that perspective that it would become necessary and essential to further evaluate the interaction of the main areas of mathematics, statistics, and computing. In this panel, we discuss the relevance and importance of incorporating modern approaches to data analytics, specifically in conjunction with biostatistics training and research. Chiefly, we aim to discuss the importance of computing and statistical thinking, as building blocks of the ideas associated with tackling projects having real-life implications, and specifically thinking about design of experiments and surveys, data collection, data visualization, modeling, modeling interpretation, and decision-making. We review these concepts from the structural, logistical, and budgetary point of views.

204 : Deciding What to Write About in the SAS World
Jane Eslinger
Thursday, September 05, 2024, 10:00 AM – 10:30 AM, Location: Resource Central
Download Slide Presentation (PDF)

I have written multiple conference papers, demos, training seminars, hands-on workshops, and college a course on SAS programming. As seasoned as I am, it is not always easy to come up with topics or decide how to tackle the topic. I’ll lead this SASquatchen discussion on how I came up with content over the years, as well as how I think others could approach writing in the SAS world. Come with questions or your own ideas and be ready to share!