The lasso is one of the most familiar tools in sparse regression, and cyclic coordinate descent is the engine behind widely used implementations such as glmnet. The reason is simple: once every other coefficient is fixed, the remaining one-dimensional lasso problem has an exact soft-thresholding solution. Its KKT conditions then say precisely which coefficients should be zero and when the algorithm has reached the optimum.
Let $\v y \in \bb{R}^n$ be a centered response and $\m X = [\v x_1,\ldots,\v x_p] \in \bb{R}^{n\times p}$ a centered design matrix. Following the standard lasso setup, scale every column so that $\|\v x_j\|_2=1$. The linear model is
$$ \v y = \m X\v\beta + \v\epsilon, \qquad \v\epsilon \sim \c N(\v 0,\sigma^2\m I_n). $$Least squares keeps every predictor. The lasso of Tibshirani (1996) adds an $\ell_1$ penalty that sets some coefficients exactly to zero:
$$ \begin{aligned} \widehat{\v\beta} &= \argmin_{\v\beta\in\bb{R}^p} f(\v\beta),\\ f(\v\beta) &:= \frac{1}{2n}\|\v y-\m X\v\beta\|_2^2 + \lambda\|\v\beta\|_1. \end{aligned} $$Write $\alpha:=n\lambda$ and define the soft-thresholding operator
$$ S_\alpha(z) := \operatorname{sign}(z)(|z|-\alpha)_+. $$This operator is the exact solution of the scalar lasso problem
$$ \argmin_{u\in\bb R} \left\{\frac12(z-u)^2+\alpha|u|\right\} =S_\alpha(z). $$With several correlated predictors, the raw correlation $\v x_j^\top\v y$ no longer decides whether $\beta_j$ is zero. What matters is the residual left by all the other coefficients. Let $\v r(\v\beta):=\v y-\m X\v\beta$. Since the absolute value is not differentiable at zero, its coordinate-wise subgradient is
$$ \partial |\beta_j| = \begin{cases} \{\operatorname{sign}(\beta_j)\}, & \beta_j\ne0,\\[0.3em] [-1,1], & \beta_j=0. \end{cases} $$The objective is convex, so requiring $0$ to lie in its subdifferential gives a necessary and sufficient optimality certificate
$$ \v x_j^\top \v r(\widehat{\v\beta}) \in \alpha\,\partial|\widehat\beta_j|, \qquad j=1,\ldots,p. $$The residual correlation sits exactly on a threshold boundary.
The residual correlation may lie anywhere inside the threshold interval.
Coordinate descent is a general strategy for minimizing a multivariable objective $F:\bb R^p\to\bb R$. At the current point $\v\beta$, choose a coordinate $j$, keep every other coordinate fixed, and minimize only along the $j$th axis:
$$ u_j^\star \in\argmin_{u\in\bb R} F(\beta_1,\ldots,\beta_{j-1},u,\beta_{j+1},\ldots,\beta_p), \qquad \beta_j\leftarrow u_j^\star. $$The new point differs from the old one in just one entry, so the path consists of axis-aligned moves. A cyclic scheme visits $1,\ldots,p$ in order; randomized and greedy rules choose the next coordinate differently. One pass through all coordinates is called a sweep, and sweeps are repeated until the updates or an optimality measure are small.
Return now to the standardized lasso. Multiplying $f$ by $n$ does not change its minimizer and turns the penalty into $\alpha\|\v\beta\|_1$. When coordinate $j$ is selected, all other coefficients are fixed. Their contribution is removed from the response to form the partial residual
$$ \v r_{\neg j} :=\v y-\sum_{k\ne j}\v x_k\beta_k =\v r+\v x_j\beta_j. $$The remaining one-dimensional lasso problem has the soft-thresholded solution
$$ \begin{aligned} \beta_j^{\mathrm{new}} &=\argmin_{u\in\bb R} \left\{\frac12\|\v r_{\neg j}-\v x_j u\|_2^2+\alpha|u|\right\}\\ &=S_\alpha(\v x_j^\top\v r_{\neg j}) =S_\alpha(\v x_j^\top\v r+\beta_j), \qquad \|\v x_j\|_2=1. \end{aligned} $$Only one coefficient changes, so the residual is updated without recomputing $\m X\v\beta$:
$$ \v r\leftarrow\v r-\v x_j (\beta_j^{\mathrm{new}}-\beta_j^{\mathrm{old}}). $$Thus the general coordinate-descent step becomes especially cheap for the lasso: compute one residual correlation, soft-threshold it, and update the residual. Cycling these exact scalar solutions drives the full objective downward.
The example below follows coordinate descent with two correlated predictors. Each move minimizes one scalar subproblem exactly while the other coefficient stays fixed.
Left: coordinate descent traces an axis-aligned path because only one coefficient moves at a time; orange segments update $\beta_1$ and green segments update $\beta_2$. Right: the corresponding scalar objective is shown with the other coefficient fixed. The open point is the exact soft-thresholded minimizer used for the next update. The status line reports the largest remaining lasso KKT discrepancy. The slider sets the correlation $\rho$ between the two predictors while holding the solution $\widehat{\v\beta}$ fixed: as $\rho$ rises the contours stretch into a narrow diagonal valley, and the same minimizer takes many more zig-zags to reach.
Cyclic coordinate descent is reliable and often very fast. But a sweep moves one coefficient at a time. With strongly correlated predictors, information can pass slowly between coordinates, and many sweeps may be required before every KKT condition is tight.