2. Visualize: Scatter, Box, and Violin Plots

David Gerbing

lessR provides many versions of a scatter plot with its XY() function for one or two variables with an option to provide a separate scatterplot for each level of one or two categorical variables. Access all scatterplots with the same simple syntax. The first variable listed without a parameter name, the x parameter, is plotted along the x-axis. Any second variable listed without a parameter name, the y parameter, is plotted along the y-axis. Each parameter may be represented by a continuous or categorical variable, a single variable or a vector of variables.

XY() also plots time series data when the x-axis variable is a Date variable. See the Time vignette for those examples.

The Data

Illustrate with the Employee data included as part of lessR.

d <- Read("Employee")
## 
## >>> Suggestions
## Recommended binary format for data files: feather
##   Create with Write(d, "your_file", format="feather")
## More details about your data, Enter:  details()  for d, or  details(name)
## 
## Data Types
## ------------------------------------------------------------
## character: Non-numeric data values
## integer: Numeric data values, integers only
## double: Numeric data values with decimal digits
## ------------------------------------------------------------
## 
##     Variable                  Missing  Unique 
##         Name     Type  Values  Values  Values   First and last values
## ------------------------------------------------------------------------------------------
##  1     Years   integer     36       1      16   7  NA  7 ... 1  2  10
##  2    Gender character     37       0       2   M  M  W ... W  W  M
##  3      Dept character     36       1       5   ADMN  SALE  FINC ... MKTG  SALE  FINC
##  4    Salary    double     37       0      37   63788.26  104494.58 ... 66508.32  67562.36
##  5    JobSat character     35       2       3   med  low  high ... high  low  high
##  6      Plan   integer     37       0       3   1  1  2 ... 2  2  1
##  7       Pre   integer     37       0      27   82  62  90 ... 83  59  80
##  8      Post   integer     37       0      22   92  74  86 ... 90  71  87
## ------------------------------------------------------------------------------------------

As an option, lessR also supports variable labels. The labels are displayed on both the text and visualization output. Each displayed label consists of the variable name juxtaposed with the corresponding label. Create the table formatted as two columns. The first column is the variable name and the second column is the corresponding variable label. Not all variables need to be entered into the table. The table can be stored as either a csv file or an Excel file.

Read the variable label file into the l data frame, currently the only permissible name for the label file.

l <- rd("Employee_lbl")
## 
## >>> Suggestions
## Recommended binary format for data files: feather
##   Create with Write(d, "your_file", format="feather")
## More details about your data, Enter:  details()  for d, or  details(name)
## 
## Data Types
## ------------------------------------------------------------
## character: Non-numeric data values
## ------------------------------------------------------------
## 
##     Variable                  Missing  Unique 
##         Name     Type  Values  Values  Values   First and last values
## ------------------------------------------------------------------------------------------
##  1     label character      8       0       8   Time of Company Employment ... Test score on legal issues after instruction
## ------------------------------------------------------------------------------------------

Display the available labels.

l
##                                                label
## Years                     Time of Company Employment
## Gender                                  Man or Woman
## Dept                             Department Employed
## Salary                           Annual Salary (USD)
## JobSat            Satisfaction with Work Environment
## Plan             1=GoodHealth, 2=GetWell, 3=BestCare
## Pre    Test score on legal issues before instruction
## Post    Test score on legal issues after instruction

Continuous Variables

Two Variables

A typical scatterplot visualizes the relationship of two continuous variables, here Years worked at a company, and annual Salary. Following is the function call to XY() for the default visualization.

Because d is the default name of the data frame that contains the variables for analysis, the data parameter that names the input data frame need not be specified. That is, no need to specify data=d, though this parameter can be explicitly included in the function call if desired.

XY(Years, Salary)
## [Interactive chart from the Plotly R package (Sievert, 2020)]

## 
## >>> Suggestions  or  enter: style(suggest=FALSE)
## XY(Years, Salary, enhance=TRUE)  # many options
## XY(Years, Salary, fill="skyblue")  # interior fill color of points
## XY(Years, Salary, fit="lm", fit_se=c(.90,.99))  # fit line, stnd errors
## XY(Years, Salary, MD_cut=6)  # Mahalanobis distance from center > 6 is an outlier 
## 
## 
## >>> Pearson's product-moment correlation 
##  
## Years: Time of Company Employment 
## Salary: Annual Salary (USD) 
##  
## Number of paired values with neither missing, n = 36 
## Sample Correlation of Years and Salary: r = 0.852 
##   
## Hypothesis Test of 0 Correlation:  t = 9.501,  df = 34,  p-value = 0.000 
## 95% Confidence Interval for Correlation:  0.727 to 0.923 
## 

Enhance the default scatterplot with parameter enhance. The visualization includes the mean of each variable indicated by the respective line through the scatterplot, the 95% confidence ellipse, labeled outliers, least-squares regression line with 95% confidence interval, and the corresponding regression line with the outliers removed.

XY(Years, Salary, enhance=TRUE)
## [Ellipse with Murdoch and Chow's function ellipse from their ellipse package]
## 
## [Interactive chart from the Plotly R package (Sievert, 2020)]

## 
## 
## >>> Suggestions  or  enter: style(suggest=FALSE)
## XY(Years, Salary, color="red")  # exterior edge color of points
## XY(Years, Salary, fit="lm", fit_se=c(.90,.99))  # fit line, stnd errors
## XY(Years, Salary, MD_cut=6)  # Mahalanobis distance from center > 6 is an outlier 
## 
## >>> Outlier analysis with squared Mahalanobis Distance 
##  
##   MD                   ID 
## -----                ----- 
## 8.34      Correll, Trevon 
## 7.73        Capelle, Adam 
##  
## 5.83   Korhalkar, Jessica 
## 5.77        James, Leslie 
## 3.92          Hoang, Binh 
## ...                   ... 
## 
## 
## >>> Pearson's product-moment correlation 
##  
## Years: Time of Company Employment 
## Salary: Annual Salary (USD) 
##  
## Number of paired values with neither missing, n = 36 
## Sample Correlation of Years and Salary: r = 0.852 
##   
## Hypothesis Test of 0 Correlation:  t = 9.501,  df = 34,  p-value = 0.000 
## 95% Confidence Interval for Correlation:  0.727 to 0.923 
## 

The default for formatting both axis labels is to round numeric values of thousands, such as 100000 to 100K. With parameter axis_fmt, this default of to {"K"} can be changed. Also can specify {","} to insert commas in large numbers with a decimal point or {"."} to insert periods, or {""} to turn off formatting. The value of {"K"} can also be combined with {","} or {"."} by forming a vector of values, such as c("K", ",").

Axis labels can also be formatted by adding a prefix to a numeric value with the parameters axis_x_prefix and axis_y_prefix, such as $ or €. The specified value can be multiple characters, such as for the Brazilian currency, R$.

XY(Years, Salary, axis_fmt=",", axis_y_prefix="$")
## [Interactive chart from the Plotly R package (Sievert, 2020)]

## 
## >>> Suggestions  or  enter: style(suggest=FALSE)
## XY(Years, Salary, enhance=TRUE)  # many options
## XY(Years, Salary, fill="skyblue")  # interior fill color of points
## XY(Years, Salary, fit="lm", fit_se=c(.90,.99))  # fit line, stnd errors
## XY(Years, Salary, MD_cut=6)  # Mahalanobis distance from center > 6 is an outlier 
## 
## 
## >>> Pearson's product-moment correlation 
##  
## Years: Time of Company Employment 
## Salary: Annual Salary (USD) 
##  
## Number of paired values with neither missing, n = 36 
## Sample Correlation of Years and Salary: r = 0.852 
##   
## Hypothesis Test of 0 Correlation:  t = 9.501,  df = 34,  p-value = 0.000 
## 95% Confidence Interval for Correlation:  0.727 to 0.923 
## 

A variety of fit lines can be plotted. The available values: "loess" for general non-linear fit, "lm" for linear least squares, "null" for the null (flat line) model, "exp" for the exponential growth and decay, "quad" for the quadratic model, and power for the general power beyond 2. Setting fit to TRUE plots the "loess" line. With the value of power, specify the value of the root with parameter fit_power.

Here, plot the general non-linear fit. For emphasis set fit_errors to TRUE to plot the residuals from the line. The sum of the squared errors is displayed to facilitate the comparison of different models.

XY(Years, Salary, fit="loess", fit_errors=TRUE)
## [Interactive chart from the Plotly R package (Sievert, 2020)]

## 
## 
## >>> Suggestions  or  enter: style(suggest=FALSE)
## XY(Years, Salary, enhance=TRUE)  # many options
## XY(Years, Salary, fill="skyblue")  # interior fill color of points
## XY(Years, Salary, MD_cut=6)  # Mahalanobis distance from center > 6 is an outlier 
## 
##    Loess Model MSE = 100,834,065.368
## 

Next, plot the exponential fit and show the residuals from the exponential curve. These data are approximately linear so the exponential curve does not vary far from a straight line. The function displays the corresponding sum of squared errors to assist in comparing various models to each other.

XY(Years, Salary, fit="exp", fit_errors=TRUE)
## [Interactive chart from the Plotly R package (Sievert, 2020)]

## 
## 
## >>> Suggestions  or  enter: style(suggest=FALSE)
## XY(Years, Salary, enhance=TRUE)  # many options
## XY(Years, Salary, fill="skyblue")  # interior fill color of points
## XY(Years, Salary, MD_cut=6)  # Mahalanobis distance from center > 6 is an outlier 
## 
## Regressed linearized data of transformed data values of Salary with log() 
##   Line: b0 = 10.959    b1 = 0.036
##   Linear Model MSE = 0.0168   Rsq = 0.725
##  
## Fit to the data with back transform exp() of linear regression model 
##  Model MSE = 127,930,074.484 
## 
## 

The parameter transforms the y variable to the specified power from the default of 1 before doing the regression analysis. The availability of this parameter provides for a wide range of modifications to the underlying functional form of the fit curve.

Three Variables

Map a continuous variable, such as Pre, to the plotted points with the pt_size parameter, a bubble plot.

XY(Years, Salary, pt_size=Pre)
## [Interactive chart from the Plotly R package (Sievert, 2020)]

##  
##  
## 
## Some Parameter values (can be manually set) 
## ------------------------------------------------------- 
## radius: 0.12    size of largest bubble 
## power: 0.50     relative bubble sizes

Indicate multiple variables to plot along either axis with a vector defined according to the base R function c(). Plot the linear model for each variable according to the fit parameter set to "lm". By default, when multiple lines are plotted on the same panel, the confidence interval is turned off by internally setting the parameter fit_se set to 0. Explicitly override this parameter value as needed.

XY(c(Pre, Post), Salary, fit="lm", fit_se=0)

## 
## 
## >>> Suggestions  or  enter: style(suggest=FALSE)
## XY(c(Pre, Post), Salary, enhance=TRUE)  # many options
## XY(c(Pre, Post), Salary, fill="skyblue")  # interior fill color of points
## XY(c(Pre, Post), Salary, out_cut=.10)  # label top 10% from center as outliers 
## 
## 
## >>> Pearson's product-moment correlation 
##  
## Post: Test score on legal issues after instruction 
## Salary: Annual Salary (USD) 
##  
## Number of paired values with neither missing, n = 37 
## Sample Correlation of Post and Salary: r = -0.070 
##   
## Hypothesis Test of 0 Correlation:  t = -0.416,  df = 35,  p-value = 0.680 
## 95% Confidence Interval for Correlation:  -0.385 to 0.260 
## 

VBS Plot of 5 Variables

Read the data and convert the values of numerically valued categorical variables to meaningful labels.

d <- Read("Cars93")
## 
## >>> Suggestions
## Recommended binary format for data files: feather
##   Create with Write(d, "your_file", format="feather")
## More details about your data, Enter:  details()  for d, or  details(name)
## 
## Data Types
## ------------------------------------------------------------
## character: Non-numeric data values
## integer: Numeric data values, integers only
## double: Numeric data values with decimal digits
## ------------------------------------------------------------
## 
##       Variable                  Missing  Unique 
##           Name     Type  Values  Values  Values   First and last values
## ------------------------------------------------------------------------------------------
##  1        Make character     93       0      32   Acura  Acura ... Volvo  Volvo
##  2        Type character     93       0       6   Small  Midsize ... Compact  Midsize
##  3    MinPrice    double     93       0      79   12.9  29.2  25.9 ... 22.9  21.8  24.8
##  4    MidPrice    double     93       0      81   15.9  33.9  29.1 ... 23.3  22.7  26.7
##  5    MaxPrice    double     93       0      79   18.8  38.7  32.3 ... 23.7  23.5  28.5
##  6     MPGcity   integer     93       0      21   25  18  20 ... 18  21  20
##  7    MPGhiway   integer     93       0      22   31  25  26 ... 25  28  28
##  8     Airbags   integer     93       0       3   0  2  1 ... 0  1  2
##  9  DriveTrain   integer     93       0       3   1  1  1 ... 1  0  1
## 10   Cylinders   integer     92       1       5   4  6  6 ... 6  4  5
## 11      Engine    double     93       0      26   1.8  3.2  2.8 ... 2.8  2.3  2.4
## 12          HP   integer     93       0      57   140  200  172 ... 178  114  168
## 13         RPM   integer     93       0      24   6300  5500  5500 ... 5800  5400  6200
## 14     RevMile   integer     93       0      78   2890  2335  2280 ... 2385  2215  2310
## 15      Manual   integer     93       0       2   1  1  1 ... 1  1  1
## 16     FuelCap    double     93       0      38   13.2  18  16.9 ... 18.5  15.8  19.3
## 17     PassCap   integer     93       0       6   5  5  5 ... 4  5  5
## 18      Length   integer     93       0      51   177  195  180 ... 159  190  184
## 19   Wheelbase   integer     93       0      27   102  115  102 ... 97  104  105
## 20       Width   integer     93       0      16   68  71  67 ... 66  67  69
## 21       Uturn   integer     93       0      14   37  38  37 ... 36  37  38
## 22    RearSeat    double     91       2      24   26.5  30  28 ... 26  29.5  30
## 23      LugCap   integer     82      11      16   11  15  14 ... 15  14  15
## 24      Weight   integer     93       0      81   2705  3560  3375 ... 2810  2985  3245
## 25      Source   integer     93       0       2   0  0  0 ... 0  0  0
## ------------------------------------------------------------------------------------------
d$Airbags <- factor(d$Airbags, levels=0:2, labels=c("none", "driver", "drv+pas"))
d$DriveTrain <- factor(d$DriveTrain, levels=0:2, labels=c("rear", "front", "all"))
d$Manual <- factor(d$Manual, levels=0:1, labels=c("Not_Avail", "Available"))

Visualize the scatterplot of MPGhiway and HP, stratified against three categorical variables: Airbags plotted in different colors for each scatterplot, and separate scatterplots for all six combinations of the levels of DriveTrain and Manual.

XY(x=MPGhiway, y=HP, by=Airbags, facet=c(DriveTrain, Manual))
## [Trellis (facet) graphics from Deepayan Sarkar's lattice package]

## 
## ---------- Summary Statistics for MPGhiway
## 
## Airbags   n  Mean  Median  SD  IQR  Min  Max
##    none  34    31      30   6    6   20   50
##  driver  43    28      28   5    4   20   46
## drv+pas  16    27      28   2    2   23   31
## 
## DriveTrain   n  Mean  Median  SD  IQR  Min  Max
##       rear  16    26      26   2    3   22   30
##      front  67    30      29   5    6   21   50
##        all  10    26      24   6    9   20   37
## 
##    Manual   n  Mean  Median  SD  IQR  Min  Max
## Not_Avail  32    26      26   3    3   20   31
## Available  61    31      30   6    6   20   50

Scatterplot Matrix

To plot a scatterplot matrix, specify multiple variables for the first parameter value, x, repeated for the second parameter, y. Define these multiple variables as a vector, such as defined by c(). Request the non-linear fit line and corresponding confidence interval by specifying TRUE or loess for the fit parameter. Request a linear fit line with the value of "lm".

d <- Read("Employee")
## 
## >>> Suggestions
## Recommended binary format for data files: feather
##   Create with Write(d, "your_file", format="feather")
## More details about your data, Enter:  details()  for d, or  details(name)
## 
## Data Types
## ------------------------------------------------------------
## character: Non-numeric data values
## integer: Numeric data values, integers only
## double: Numeric data values with decimal digits
## ------------------------------------------------------------
## 
##     Variable                  Missing  Unique 
##         Name     Type  Values  Values  Values   First and last values
## ------------------------------------------------------------------------------------------
##  1     Years   integer     36       1      16   7  NA  7 ... 1  2  10
##  2    Gender character     37       0       2   M  M  W ... W  W  M
##  3      Dept character     36       1       5   ADMN  SALE  FINC ... MKTG  SALE  FINC
##  4    Salary    double     37       0      37   63788.26  104494.58 ... 66508.32  67562.36
##  5    JobSat character     35       2       3   med  low  high ... high  low  high
##  6      Plan   integer     37       0       3   1  1  2 ... 2  2  1
##  7       Pre   integer     37       0      27   82  62  90 ... 83  59  80
##  8      Post   integer     37       0      22   92  74  86 ... 90  71  87
## ------------------------------------------------------------------------------------------
XY(c(Salary, Years, Pre, Post), c(Salary, Years, Pre, Post), fit="lm")

Smoothed, Contoured, and Binned Scatterplots

Smoothing and binning are two procedures for visualizing a relationship with many data values.

To obtain a larger data set, in this example generate random data with base R rnorm(), then plot. XY() first checks the presence of the specified variables in the global environment (workspace). If not there, then from a data frame, of which the default value is d. Here, randomly generate values from normal populations for x and y in the workspace.

set.seed(13)
x=rnorm(4000)
y= 8*x + rnorm(4000,1, 30)
XY(x, y, data=NULL)
## [Interactive chart from the Plotly R package (Sievert, 2020)]

## 
## >>> Suggestions  or  enter: style(suggest=FALSE)
## XY(x, y, enhance=TRUE)  # many options
## XY(x, y, color="red")  # exterior edge color of points
## XY(x, y, fit="lm", fit_se=c(.90,.99))  # fit line, stnd errors
## XY(x, y, out_cut=.10)  # label top 10% from center as outliers 
## 
## 
## >>> Pearson's product-moment correlation 
##  
## Number of paired values with neither missing, n = 4000 
## Sample Correlation of x and y: r = 0.251 
##   
## Hypothesis Test of 0 Correlation:  t = 16.397,  df = 3998,  p-value = 0.000 
## 95% Confidence Interval for Correlation:  0.222 to 0.280 
## 

With large data sets, even for continuous variables there can be much over-plotting of points. One strategy to address this issue smooths the scatterplot by setting the form parameter to smooth. The individual points superimposed on the smoothed plot are potential outliers. The default number of plotted outliers is 100. Turn off the plotting of outliers completely by setting parameter smooth_points to 0. Show the linear trend with fit set to "lm".

XY(x, y, form="smooth", fit="lm", data=NULL)

## 
## 
## >>> Suggestions  or  enter: style(suggest=FALSE)
## XY(x, y, enhance=TRUE)  # many options
## XY(x, y, fill="skyblue")  # interior fill color of points
## XY(x, y, MD_cut=6)  # Mahalanobis distance from center > 6 is an outlier 
## 
## 
## >>> Pearson's product-moment correlation 
##  
## Number of paired values with neither missing, n = 4000 
## Sample Correlation of x and y: r = 0.251 
##   
## Hypothesis Test of 0 Correlation:  t = 16.397,  df = 3998,  p-value = 0.000 
## 95% Confidence Interval for Correlation:  0.222 to 0.280 
##   
## 
##   Line: b0 = 1.03068757    b1 = 7.91963664
##   Linear Model MSE = 917.03180812   Rsq = 0.063
## 

Another strategy for alleviating over-plotting makes the fill color mostly transparent with the transparency parameter, or turn off completely by setting fill to "off". The closer the value of trans is to 1, the more transparent is the fill.

XY(x, y, transparency=0.95, data=NULL)
## [Interactive chart from the Plotly R package (Sievert, 2020)]

## 
## >>> Suggestions  or  enter: style(suggest=FALSE)
## XY(x, y, enhance=TRUE)  # many options
## XY(x, y, color="red")  # exterior edge color of points
## XY(x, y, fit="lm", fit_se=c(.90,.99))  # fit line, stnd errors
## XY(x, y, MD_cut=6)  # Mahalanobis distance from center > 6 is an outlier 
## 
## 
## >>> Pearson's product-moment correlation 
##  
## Number of paired values with neither missing, n = 4000 
## Sample Correlation of x and y: r = 0.251 
##   
## Hypothesis Test of 0 Correlation:  t = 16.397,  df = 3998,  p-value = 0.000 
## 95% Confidence Interval for Correlation:  0.222 to 0.280 
## 

Contour plots are another effective way to visualize scatter plots with much data. The parameter contour_n sets the number of filled density bands, 20 by default. If there are extreme outliers, the axes extend to their maximum and minimum values, typically leaving much white space around the visible contour plot; the extreme values of outlier points with low density round down to zero on the color scale. The parameters pad_x and pad_y, each with a default value of c(0,0), pad the plot at the low and high end of an axis. Increase a value to add more space on that side.

XY(x, y, form="contour", data=NULL)

## 
## >>> Suggestions  or  enter: style(suggest=FALSE)
## XY(x, y, enhance=TRUE)  # many options
## XY(x, y, color="red")  # exterior edge color of points
## XY(x, y, fit="lm", fit_se=c(.90,.99))  # fit line, stnd errors
## XY(x, y, out_cut=.10)  # label top 10% from center as outliers 
## 
## 
## >>> Pearson's product-moment correlation 
##  
## Number of paired values with neither missing, n = 4000 
## Sample Correlation of x and y: r = 0.251 
##   
## Hypothesis Test of 0 Correlation:  t = 16.397,  df = 3998,  p-value = 0.000 
## 95% Confidence Interval for Correlation:  0.222 to 0.280 
## 

The density forms also render the cells of a scatterplot matrix, which is where over-plotting is worst of all: every cell holds the same many observations. Each cell estimates its own density over its own pair of variables, so nothing is pooled across cells, and the correlations remain in the upper triangle. A cell defaults to 8 bands rather than the 20 of a single plot, too fine to resolve at cell size. The matrix reads its variables from a data frame rather than from the workspace, so collect them into one first.

dm <- data.frame(a=x, b=y, g=20*x + rnorm(4000, 0, 10),
                 h=rnorm(4000))
XY(c(a,b,g,h), c(a,b,g,h), form="contour", data=dm)

Another way to visualize a relationship when there are many data points is to bin the x-axis. Specify the number of bins with parameter n_bins. XY() then computes the mean of y for each bin and connects the means by line segments. This procedure plots the conditional means by default without any assumption of form such as linearity. Specify the stat parameter for median to compute the median of y for each bin. The standard XY() parameters fill, color, pt_size and segments also apply.

XY(x, y, n_bins=5, data=NULL)
## [Interactive chart from the Plotly R package (Sievert, 2020)]

## 
## Table: Summary Stats 
##  
##                x          y 
## -------  -------  --------- 
## n           4000       4000 
## n.miss         0          0 
## min       -3.239   -104.740 
## max        3.589    112.460 
## mean      -0.003      1.006 
## 
##  
## Table: mean of y for levels of x 
##  
##                   bin      n    midpt      mean 
## ---  ----------------  -----  -------  -------- 
## 1     [-3.246,-1.873]    116   -2.560   -16.734 
## 2     (-1.873,-0.508]   1090   -1.191    -5.699 
## 3      (-0.508,0.858]   2001    0.175     0.848 
## 4       (0.858,2.223]    743    1.541    12.374 
## 5       (2.223,3.596]     50    2.909    25.696

Each preceding strategy addresses over-plotting that arises from density, many points crowded into a region of continuous space. Over-plotting also arises from exact coincidence, when discrete or rounded values place observations at identical coordinates. For that situation set form to "sunflower". Each plotted point carries one petal for every observation at its coordinate, so a plain dot marks a single observation and a sunflower marks several. Here Years and Pre are both integer valued, so coordinates repeat.

XY(Years, Pre, form="sunflower")

Petals appear only where observations coincide. For continuous data such as Years and Salary, where no two employees share a salary to the cent, every point plots as a plain dot.

Stratify the sunflower with facet, one panel per group. The larger Cars93 data set provides more coincident coordinates than the 37 rows of Employee. Horsepower and city gas mileage are both recorded as integers, so cars repeat each other’s coordinates.

d <- Read("Cars93")
## 
## >>> Suggestions
## Recommended binary format for data files: feather
##   Create with Write(d, "your_file", format="feather")
## More details about your data, Enter:  details()  for d, or  details(name)
## 
## Data Types
## ------------------------------------------------------------
## character: Non-numeric data values
## integer: Numeric data values, integers only
## double: Numeric data values with decimal digits
## ------------------------------------------------------------
## 
##       Variable                  Missing  Unique 
##           Name     Type  Values  Values  Values   First and last values
## ------------------------------------------------------------------------------------------
##  1        Make character     93       0      32   Acura  Acura ... Volvo  Volvo
##  2        Type character     93       0       6   Small  Midsize ... Compact  Midsize
##  3    MinPrice    double     93       0      79   12.9  29.2  25.9 ... 22.9  21.8  24.8
##  4    MidPrice    double     93       0      81   15.9  33.9  29.1 ... 23.3  22.7  26.7
##  5    MaxPrice    double     93       0      79   18.8  38.7  32.3 ... 23.7  23.5  28.5
##  6     MPGcity   integer     93       0      21   25  18  20 ... 18  21  20
##  7    MPGhiway   integer     93       0      22   31  25  26 ... 25  28  28
##  8     Airbags   integer     93       0       3   0  2  1 ... 0  1  2
##  9  DriveTrain   integer     93       0       3   1  1  1 ... 1  0  1
## 10   Cylinders   integer     92       1       5   4  6  6 ... 6  4  5
## 11      Engine    double     93       0      26   1.8  3.2  2.8 ... 2.8  2.3  2.4
## 12          HP   integer     93       0      57   140  200  172 ... 178  114  168
## 13         RPM   integer     93       0      24   6300  5500  5500 ... 5800  5400  6200
## 14     RevMile   integer     93       0      78   2890  2335  2280 ... 2385  2215  2310
## 15      Manual   integer     93       0       2   1  1  1 ... 1  1  1
## 16     FuelCap    double     93       0      38   13.2  18  16.9 ... 18.5  15.8  19.3
## 17     PassCap   integer     93       0       6   5  5  5 ... 4  5  5
## 18      Length   integer     93       0      51   177  195  180 ... 159  190  184
## 19   Wheelbase   integer     93       0      27   102  115  102 ... 97  104  105
## 20       Width   integer     93       0      16   68  71  67 ... 66  67  69
## 21       Uturn   integer     93       0      14   37  38  37 ... 36  37  38
## 22    RearSeat    double     91       2      24   26.5  30  28 ... 26  29.5  30
## 23      LugCap   integer     82      11      16   11  15  14 ... 15  14  15
## 24      Weight   integer     93       0      81   2705  3560  3375 ... 2810  2985  3245
## 25      Source   integer     93       0       2   0  0  0 ... 0  0  0
## ------------------------------------------------------------------------------------------
XY(HP, MPGcity, form="sunflower", facet=Source)

Each panel has its own set of coordinates, so each has its own petal counts. The by parameter does not apply to this form, because overlaid groups would interleave their petals at a shared coordinate, leaving no count to read. To stratify within a single panel, use form="hexbin".

Return to the Employee data for the sections that follow.

d <- Read("Employee", quiet=TRUE)

Quantile-Quantile Chart

Every scatterplot so far reads both of its variables from the data. A quantile-quantile chart reads one and generates the other. Name the theoretical distribution in the x position with the keyword .normal, and XY() generates its quantiles from the values of y. The keyword names no column of data, in the same way that the keyword .index generates the consecutive integers of a run chart.

XY(.normal, Salary)
## [Interactive chart from the Plotly R package (Sievert, 2020)]

## 
## >>> Suggestions  or  enter: style(suggest=FALSE)
## XY(.normal, Salary, enhance=TRUE)  # many options
## XY(.normal, Salary, color="red")  # exterior edge color of points
## XY(.normal, Salary, fit="lm", fit_se=c(.90,.99))  # fit line, stnd errors
## XY(.normal, Salary, MD_cut=6)  # Mahalanobis distance from center > 6 is an outlier

The generated quantiles carry the mean and standard deviation of Salary, so both axes are in dollars and the reference against which the data are read is the 45-degree line through the origin. A point above the line is larger than a normal distribution accounts for, a point below it smaller.

Four distributions are available, each fit to y by the method of moments, so none asks for a shape parameter: .normal, .lognormal, .exponential, and .uniform. Salary trails off to the right, as the chart above shows in its upper tail, so compare it against a right-skewed distribution instead.

XY(.lognormal, Salary)
## [Interactive chart from the Plotly R package (Sievert, 2020)]

## 
## >>> Suggestions  or  enter: style(suggest=FALSE)
## XY(.lognormal, Salary, enhance=TRUE)  # many options
## XY(.lognormal, Salary, fill="skyblue")  # interior fill color of points
## XY(.lognormal, Salary, fit="lm", fit_se=c(.90,.99))  # fit line, stnd errors
## XY(.lognormal, Salary, out_cut=.10)  # label top 10% from center as outliers

The lognormal requires positive values and the exponential non-negative values. Because y accepts an expression, XY(.normal, log(Salary)) draws the same picture as the chart above; the difference is that the lognormal keyword keeps both axes in dollars rather than in log dollars.

The chart is an ordinary scatterplot, so the compositions of XY() apply. Panel the chart with facet, or overlay the groups in a single panel with by. Either way the quantiles are generated separately within each group, from that group’s own mean and standard deviation, so the one 45-degree line serves every group.

XY(.normal, Salary, facet=Gender)
## [Trellis (facet) graphics from Deepayan Sarkar's lattice package]

With by the groups share a single panel.

XY(.normal, Salary, by=Gender)
## [Interactive chart from the Plotly R package (Sievert, 2020)]

## 
## >>> Suggestions  or  enter: style(suggest=FALSE)
## XY(.normal, Salary, enhance=TRUE)  # many options
## XY(.normal, Salary, fill="skyblue")  # interior fill color of points
## XY(.normal, Salary, fit="lm", fit_se=c(.90,.99))  # fit line, stnd errors
## XY(.normal, Salary, MD_cut=6)  # Mahalanobis distance from center > 6 is an outlier
## XY(.normal, Salary, by=Gender, shape="diamond")  # diamond for points

Here the x-axis is not a single distribution against which both groups are compared. Each group is read against its own fitted normal, so a point’s x-value is the salary that group’s own fit predicts at that quantile. Because those fits carry each group’s mean and standard deviation, the two clouds separate along the reference line, which is what makes the one line correct for both: read location and spread from position along the line, and departure from normality from distance away from it. Overlaid groups can interleave, though, so facet reads more cleanly for more than two.

Set data to NULL to read y from the workspace instead of a data frame, which displays the distribution of a computed vector such as the residuals of a fitted model.

m <- lm(Salary ~ Years, data=d)
XY(.normal, residuals(m), data=NULL)
## [Interactive chart from the Plotly R package (Sievert, 2020)]

## 
## >>> Suggestions  or  enter: style(suggest=FALSE)
## XY(.normal, residuals(m), enhance=TRUE)  # many options
## XY(.normal, residuals(m), fill="skyblue")  # interior fill color of points
## XY(.normal, residuals(m), fit="lm", fit_se=c(.90,.99))  # fit line, stnd errors
## XY(.normal, residuals(m), MD_cut=6)  # Mahalanobis distance from center > 6 is an outlier

Categorical and Continuous Variables

A mixture of categorical and continuous variables can be plotted a variety of ways, as illustrated below.

Two Continuous, One Categorical

Plot a scatterplot of two continuous variables for each level of a categorical variable on the same panel with the by parameter. Here, plot Years and Salary each for the two levels of Gender in the data. Colors and geometric plot shapes can distinguish between the plots. For all variables except an ordered factor, the default plots according to the default qualitative color palette, "hues", with the geometric shape of a point.

XY(Years, Salary, by=Gender)
## [Interactive chart from the Plotly R package (Sievert, 2020)]

## 
## >>> Suggestions  or  enter: style(suggest=FALSE)
## XY(Years, Salary, enhance=TRUE)  # many options
## XY(Years, Salary, color="red")  # exterior edge color of points
## XY(Years, Salary, fit="lm", fit_se=c(.90,.99))  # fit line, stnd errors
## XY(Years, Salary, out_cut=.10)  # label top 10% from center as outliers
## XY(Years, Salary, by=Gender, shape="diamond")  # diamond for points

Change the plot colors with the fill (interior) and color (exterior or edge) parameters. Because there are two levels of the by variable, specify two fill colors and two edge colors each with an R vector defined by the c() function. Also, include the regression line for each group with the fit parameter and increase the size of the plotted points with the pt_size parameter.

XY(Years, Salary, by=Gender, pt_size=2, fit="lm",
     fill=c(M="olivedrab3", W="gold1"), 
     color=c(M="darkgreen", W="gold4")
)
## [Interactive chart from the Plotly R package (Sievert, 2020)]

## 
## 
## >>> Suggestions  or  enter: style(suggest=FALSE)
## XY(Years, Salary, enhance=TRUE)  # many options
## XY(Years, Salary, out_cut=.10)  # label top 10% from center as outliers
## XY(Years, Salary, by=Gender, fill=c(M="olivedrab3", W="gold1"), color=c(M="darkgreen", W="gold4"), pt_size=2, fit="lm", shape="diamond")  # diamond for points 
## 
## Gender: M  Line: b0 = 40842.335    b1 = 4047.307
##   Linear Model MSE = 107,647,877.258   Rsq = 0.819
##  
## Gender: W  Line: b0 = 57109.787    b1 = 2882.272
##   Linear Model MSE = 144,700,624.695   Rsq = 0.598
## 

Change the plotted shapes with the pt_shape parameter. The default value is "circle" with both an exterior color and filled interior, specified with "color" and "fill". Other possible values, with fillable interiors, are "circle", "square", "diamond", "triup" (triangle up), and "tridown" (triangle down). Other possible values include all uppercase and lowercase letters, all digits, and most punctuation characters. The numbers 0 through 25 defined by the R points() function also apply. If plotting levels according to by, then list one shape for each level to be plotted.

Or, request default shapes across the different by groups by setting parameter shapes to "vary".

XY(Years, Salary, by=Gender, pt_shape="vary")
## [Interactive chart from the Plotly R package (Sievert, 2020)]

## 
## >>> Suggestions  or  enter: style(suggest=FALSE)
## XY(Years, Salary, enhance=TRUE)  # many options
## XY(Years, Salary, color="red")  # exterior edge color of points
## XY(Years, Salary, fit="lm", fit_se=c(.90,.99))  # fit line, stnd errors
## XY(Years, Salary, out_cut=.10)  # label top 10% from center as outliers

A Trellis (facet) plot creates a separate panel for the plot of each level of the categorical variable. Generate Trellis plots with the facet parameter. In this example, plot the best-fit linear model for the data in each panel according to the fit parameter. By default, the 95% confidence interval for each line is also displayed.

XY(Years, Salary, facet=Gender, fit="lm")
## [Trellis (facet) graphics from Deepayan Sarkar's lattice package]

## 
## Regression analysis of linearized Salary values
## Need back transformation of regression model to compute predicted values
## 
## Gender 1  Line: b0 = 40842.335  b1 = 4047.307   Fit: MSE =    Rsq = 0.819
## 
## Gender 2  Line: b0 = 57109.787  b1 = 2882.272   Fit: MSE =    Rsq = 0.598
## 
## ---------- Summary Statistics for Years
## 
## Gender   n  Mean  Median  SD  IQR  Min  Max
##      M  17    12      13   5    5    5   24
##      W  19     7       6   5    6    1   18

Turn off the confidence interval by setting the parameter fit_se to 0 for the value of the confidence level.

Two Continuous, Three Categorical

To illustrate, first, the data. Use the Cars93 data set that is installed with lessR, which describes characteristics of 1993 cars.

d <- Read("Cars93")
## 
## >>> Suggestions
## Recommended binary format for data files: feather
##   Create with Write(d, "your_file", format="feather")
## More details about your data, Enter:  details()  for d, or  details(name)
## 
## Data Types
## ------------------------------------------------------------
## character: Non-numeric data values
## integer: Numeric data values, integers only
## double: Numeric data values with decimal digits
## ------------------------------------------------------------
## 
##       Variable                  Missing  Unique 
##           Name     Type  Values  Values  Values   First and last values
## ------------------------------------------------------------------------------------------
##  1        Make character     93       0      32   Acura  Acura ... Volvo  Volvo
##  2        Type character     93       0       6   Small  Midsize ... Compact  Midsize
##  3    MinPrice    double     93       0      79   12.9  29.2  25.9 ... 22.9  21.8  24.8
##  4    MidPrice    double     93       0      81   15.9  33.9  29.1 ... 23.3  22.7  26.7
##  5    MaxPrice    double     93       0      79   18.8  38.7  32.3 ... 23.7  23.5  28.5
##  6     MPGcity   integer     93       0      21   25  18  20 ... 18  21  20
##  7    MPGhiway   integer     93       0      22   31  25  26 ... 25  28  28
##  8     Airbags   integer     93       0       3   0  2  1 ... 0  1  2
##  9  DriveTrain   integer     93       0       3   1  1  1 ... 1  0  1
## 10   Cylinders   integer     92       1       5   4  6  6 ... 6  4  5
## 11      Engine    double     93       0      26   1.8  3.2  2.8 ... 2.8  2.3  2.4
## 12          HP   integer     93       0      57   140  200  172 ... 178  114  168
## 13         RPM   integer     93       0      24   6300  5500  5500 ... 5800  5400  6200
## 14     RevMile   integer     93       0      78   2890  2335  2280 ... 2385  2215  2310
## 15      Manual   integer     93       0       2   1  1  1 ... 1  1  1
## 16     FuelCap    double     93       0      38   13.2  18  16.9 ... 18.5  15.8  19.3
## 17     PassCap   integer     93       0       6   5  5  5 ... 4  5  5
## 18      Length   integer     93       0      51   177  195  180 ... 159  190  184
## 19   Wheelbase   integer     93       0      27   102  115  102 ... 97  104  105
## 20       Width   integer     93       0      16   68  71  67 ... 66  67  69
## 21       Uturn   integer     93       0      14   37  38  37 ... 36  37  38
## 22    RearSeat    double     91       2      24   26.5  30  28 ... 26  29.5  30
## 23      LugCap   integer     82      11      16   11  15  14 ... 15  14  15
## 24      Weight   integer     93       0      81   2705  3560  3375 ... 2810  2985  3245
## 25      Source   integer     93       0       2   0  0  0 ... 0  0  0
## ------------------------------------------------------------------------------------------

Two of the categorical variables are integer coded 0 and 1, so recode to R factors to obtain more descriptive labels. For clarity, convert the relevant categorical variables to factors, including Cylinders the number of cylinders for a car, for consistency.

d$Trans <- factor(d$Manual, levels=0:1, labels=c("Auto", "Manual"))
d$Source <- factor(d$Source, levels=0:1, labels=c("Foreign", "Domestic"))
d$Cylinders <- factor(d$Cylinders, levels=c(4,6,8))

Two Continuous, Three Categorical

XY() can display the relationships for up to five variables. The two primary variables, x and y, that form the basis of the scatter plot, are continuous. Usually these two variables are listed first in the function call and so do not need their parameter names specified. Indicate two categorical variables that form the Trellis panels with parameter facet. Call these two variables the Trellis variables, which define a Trellis panel for each combination of their values. Finally, there can be a categorical grouping variable, the by variable, which plots different groups within each Trellis panel.

Plot MPGcity according to Weight. Specify the number of Cylinders and Manual transmission or not as Trellis conditioning variables to form the Trellis plot. Specify the Source of the vehicle, Foreign or Domestic as a grouping variable to plot with separate colors on each panel. Use the parameter value n_axis_x_skip=2 to include only every third axis tick label due to the lack of room to avoid overlapping labels.

XY(Weight, MPGcity, by=Source, facet=c(Cylinders,Trans), n_axis_x_skip=2, n_row=2)
## [Trellis (facet) graphics from Deepayan Sarkar's lattice package]

## 
## ---------- Summary Statistics for Weight
## 
##   Source   n  Mean  Median   SD  IQR   Min   Max
##  Foreign  39  2990    2970  538  940  2055  4100
## Domestic  48  3195    3282  565  934  1845  4105
## 
## Cylinders   n  Mean  Median   SD  IQR   Min   Max
##         4  49  2710    2705  373  520  1845  3785
##         6  31  3559    3515  265  255  2810  4105
##         8   7  3836    3935  244  210  3380  4055
## 
##  Trans   n  Mean  Median   SD  IQR   Min   Max
##   Auto  32  3569    3590  357  345  2880  4105
## Manual  55  2832    2785  471  585  1845  3805

From the visualization the patterns emerge. As Weight increases city MPG decreases. Domestic cars tend to weigh more. Foreign cars tend to have fewer cylinders, which also leads to better fuel mileage.

Categorical Variables

To avoid over-plotting, the plot of two categorical variables results in a bubble plot of their joint frequencies.

d <- Read("Employee", quiet=TRUE)
Chart(Dept, by=Gender, form="bubble")
## >>> Suggestions  or  enter: style(suggest=FALSE)
## Chart(Dept, by=Gender, form="radar")  # Plotly radar chart
## Chart(Dept, by=Gender, form="treemap")  # Plotly treemap chart
## Chart(Dept, by=Gender, form="pie")  # Plotly pie/sunburst chart
## Chart(Dept, by=Gender, form="icicle")  # Plotly icicle chart
## Chart(Dept, by=Gender, form="dot")  # Plotly dot chart
## Chart(Dept, by=Gender, form="profile")  # points connected across categories
##  
## 
## Joint and Marginal Frequencies 
## ------------------------------ 
##  
##        Dept 
## Gender   ACCT ADMN FINC MKTG SALE Sum 
##   M         2    2    3    1   10  18 
##   W         3    4    1    5    5  18 
##   Sum       5    6    4    6   15  36 
## 
## Cramer's V: 0.415 
##  
## Chi-square Test of Independence:
##      Chisq = 6.200, df = 4, p-value = 0.185 
## >>> Low cell expected frequencies, chi-squared approximation may not be accurate 
## 
## [Interactive chart from the Plotly R package (Sievert, 2020)]

The parameter radius scales the size of the bubbles according to the size of the largest displayed bubble in inches. The power parameter sets the relative size of the bubbles. The default power value of 0.5 scales the bubbles so that the area of each bubble is the value of the corresponding sizing variable. A value of 1 scales so the radius of each bubble is the value of the sizing variable, increasing the discrepancy of size between the variables.

In this example, increase the absolute size of the bubbles as well as the relative discrepancy in their sizes. If the bubbles become too large, so that the largest bubbles become truncated, increase the spacing of the respective axes with the pad_x and/or pad_y parameters.

Chart(Dept, by=Gender, radius=.55, power=0.8)
## >>> Suggestions  or  enter: style(suggest=FALSE)
## Chart(Dept, by=Gender, form="radar")  # Plotly radar chart
## Chart(Dept, by=Gender, form="treemap")  # Plotly treemap chart
## Chart(Dept, by=Gender, form="pie")  # Plotly pie/sunburst chart
## Chart(Dept, by=Gender, form="icicle")  # Plotly icicle chart
## Chart(Dept, by=Gender, form="bubble")  # Plotly bubble chart
## Chart(Dept, by=Gender, form="dot")  # Plotly dot chart
## Chart(Dept, by=Gender, form="profile")  # points connected across categories
##  
## [Interactive chart from the Plotly R package (Sievert, 2020)]

## Dept: Department Employed 
##   - by levels of - 
## Gender: Man or Woman 
## 
## Joint and Marginal Frequencies 
## ------------------------------ 
##  
##        Dept 
## Gender   ACCT ADMN FINC MKTG SALE Sum 
##   M         2    2    3    1   10  18 
##   W         3    4    1    5    5  18 
##   Sum       5    6    4    6   15  36 
## 
## Cramer's V: 0.415 
##  
## Chi-square Test of Independence:
##      Chisq = 6.200, df = 4, p-value = 0.185 
## >>> Low cell expected frequencies, chi-squared approximation may not be accurate

Input Interactive Plots

An interactive visualization lets the user in real time change parameter values to change characteristics of the visualization. To create an interactive two-variable scatterplot of continuous variables with the employee data that displays the corresponding parameters, run the function interact() with "XY" specified.

interact("XY")

To create an interactive Trellis plot as a combined violin, box, and scatter plot with the five values of Dept from the Employee data set that displays the corresponding parameters, run the function interact() with "Trellis" specified.

interact("Trellis")

The functions are not run here because interactivity requires to run directly from the R console.

Full Manual

Use the base R help() function to view the full manual for XY(). Simply enter a question mark followed by the name of the function.

?XY

More

More on Scatterplots, Time Series plots, and other visualizations from lessR and other packages such as ggplot2 at:

Gerbing, D., R Visualizations: Derive Meaning from Data, CRC Press, May, 2020, ISBN 978-1138599635.