The central design choice in an MM algorithm is the surrogate. A valid majorizer guarantees descent, but the speed of convergence depends strongly on how tightly it follows the objective near the current iterate. For $g(\v\gamma)$, the coordinates are coupled inside a matrix inverse. A matrix Cauchy-Schwarz inequality gives a tangent upper bound that separates into $p$ scalar problems with closed-form updates. The construction also introduces a simplex weight vector $\v w$: each choice of $\v w$ gives a different valid surrogate.
Titu's lemma, also known as Sedrakyan's inequality, holds for real \(\widetilde\sigma_j\) and strictly positive \(\sigma_j\):
$$ \frac{(\sum_j\widetilde\sigma_j)^2}{\sum_j\sigma_j} \le \sum_j\frac{\widetilde\sigma_j^2}{\sigma_j}. $$It is Cauchy-Schwarz in disguise. Apply $(\sum_j a_jb_j)^2\le(\sum_j a_j^2)(\sum_j b_j^2)$ with $a_j=\widetilde\sigma_j/\sqrt{\sigma_j}$ and $b_j=\sqrt{\sigma_j}$:
$$ \Bigl(\sum_j\widetilde\sigma_j\Bigr)^2 =\Bigl(\sum_j\frac{\widetilde\sigma_j}{\sqrt{\sigma_j}}\cdot\sqrt{\sigma_j}\Bigr)^2 \le\Bigl(\sum_j\frac{\widetilde\sigma_j^2}{\sigma_j}\Bigr) \Bigl(\sum_j\sigma_j\Bigr), $$then divide both sides by $\sum_j\sigma_j>0$.
For the matrix analogue, let $\m V_1,\ldots,\m V_m$ be symmetric positive-semidefinite matrices, take $\v\sigma\in\bb R_{++}^m$, and take $\widetilde{\v\sigma}\in\bb R_+^m\setminus\{\v0\}$. In addition, require both weighted sums $\sum_j\sigma_j\m V_j$ and $\sum_j\widetilde\sigma_j\m V_j$ to be positive definite, so that the ordinary inverses below exist. Then
$$ \left(\sum_j\sigma_j\m V_j\right)^{-1} \preceq \left(\sum_j\widetilde\sigma_j\m V_j\right)^{-1} \left(\sum_j\frac{\widetilde\sigma_j^2}{\sigma_j}\m V_j\right) \left(\sum_j\widetilde\sigma_j\m V_j\right)^{-1}. $$Equality holds when $\sigma_j=\widetilde\sigma_j$, which is exactly the tangency an MM surrogate needs.
Let $\v w\ge\v0$ satisfy $\v1^\top\v w=1$. Then $\m I=\sum_j w_j\m I$, and adding and subtracting $w_j\v x_j\v x_j^\top$ splits the inverse of $\m G$ into $2p$ terms:
$$ \m G^{-1}(\v\gamma) =\sum_{j=1}^p(\gamma_j+w_j)\v x_j\v x_j^\top +\sum_{j=1}^p w_j(\m I-\v x_j\v x_j^\top). $$Because each standardized predictor has $\|\v x_j\|_2=1$, every $\m I-\v x_j\v x_j^\top$ is a projection and hence positive semidefinite, so the matrix inequality applies with these $2p$ matrices as the $\m V$'s; terms whose coefficient is zero are simply dropped. Apply it at the current iterate $\widetilde{\v\gamma}$. With $\widetilde\tau_j:=|\v x_j^\top\m G(\widetilde{\v\gamma})\v y|$ and $\widetilde r:=\|\m G(\widetilde{\v\gamma})\v y\|_2$, the resulting surrogate is
$$ h_{\v w}(\v\gamma,\widetilde{\v\gamma}) =\sum_{j=1}^p\left[ \frac{(\widetilde\gamma_j+w_j)^2\widetilde\tau_j^2}{\gamma_j+w_j} +w_j(\widetilde r^2-\widetilde\tau_j^2) +\alpha^2\gamma_j \right]. $$The variables now separate. Minimizing the bound over each $\gamma_j\ge0$ gives
$$ \gamma_j^{(t+1)} =\frac{1}{\alpha} S_{\alpha w_j}\!\left((\gamma_j^{(t)}+w_j)\tau_j^{(t)}\right), \qquad j=1,\ldots,p. $$One solve for $\m G(\v\gamma^{(t)})\v y$ supplies every $\tau_j^{(t)}$, and all $p$ coordinates then update together. If $w_j=0$, the rule becomes multiplicative:
$$ \gamma_j^{(t+1)} =\gamma_j^{(t)}\frac{\tau_j^{(t)}}{\alpha}. $$A multiplicative step cannot move a zero coordinate away from zero. Positive weight is the scarce resource that lets a coefficient enter or leave the active set.
The scalar update makes this concrete. With $\alpha=1$, the plot below shows $\gamma_j^{(t+1)}$ as a function of the weight $w_j$ for two coordinates that see the same residual signal $\tau_j$: one currently zero, $\gamma_j^{(t)}=0$, and one currently active at $\gamma_j^{(t)}=0.4$. Expanding the soft-threshold,
$$ \gamma_j^{(t+1)} =\frac{1}{\alpha}\Bigl[\gamma_j^{(t)}\tau_j+w_j(\tau_j-\alpha)\Bigr]_+ . $$The bracket is affine in $w_j$ with slope $\tau_j-\alpha=-\xi_j$, where $\xi_j$ is the signed KKT gap of Part 2. Slide $\tau_j$ across the threshold: below it, weight can only shrink an active coordinate, and enough weight switches it off; above it, any positive weight activates the zero coordinate, and more weight moves both further.
Blue: a coordinate that is currently zero. Green: a coordinate currently active at $\gamma_j^{(t)}=0.4$. The dashed horizontal marks the active coordinate's current value. At $w_j=0$ both lines start at the multiplicative update $\gamma_j^{(t)}\tau_j/\alpha$, so the zero coordinate is stuck there. When $\tau_j<\alpha$ the green line reaches zero at $w_j=\gamma_j^{(t)}\tau_j/(\alpha-\tau_j)$, the open circle, which Part 5 names the deactivation threshold $c_j$.
Valid does not mean equally useful. The optimal weights minimize the surrogate's one-step upper bound over the simplex. Part 5 derives that weight problem, interprets its KKT conditions, and shows how to compute its exact solution.