Math 202A: Lecture 16

This lecture is the most important lecture in Math 202A: we will obtain a remarkably transparent of morphisms in \mathbf{FHil}, in total generality. It is perhaps surprising that this can be done; it is perhaps not surprising that it takes a couple of tries to get it right.

Take any two objects V,W and any morphism A \in \mathrm{Hom}(V,W) between them. Our job is to understand A without making any special assumptions on its structure. A natural strategy – perhaps the only strategy in such generality – is to find a canonical (in this case, meaning defined in the same way for all spaces) subspace V_0 of V on which we can understand the behavior of A completely. The hope is that we can bootstrap this local understanding of A on the special subspace V_0 to a global understanding on A on all of V.

This is yet another opportunity to point out that the quantum category is better than the classical one. Indeed, suppose we can find a subspace V_0 of V on which we can calculate the outputs of A exactly. If we think set theoretically, then in order to complete our understanding of A we would have to find a way to understand it on the set-theoretic complement

V_0^c = \{v \in V \colon v \not\in V_0\}.

But linear algebraically, in order to complete our understanding of A it is sufficient to bootstrap our understanding to the orthogonal complement

V_0^\perp=\{v \in V \colon \langle v,v_0 \rangle=0 \text{ for all }v \in V_0\}.

For example in three dimensions, this is the difference between extending from the line of understanding to the whole space with that line deleted, and extending only to the plane orthogonal to the line of understanding. The latter is a dimension reduction, the former is not.

There is one obvious candidate for an “exactly solvable” subspace V_0 canonically associated to any A \in \mathrm{Hom}(V,W) namely V_0 = \mathrm{Ker}(A). We know exactly how to compute Av for every v \in V_0, the output is Av=0_W. Unfortunately, the kernel is not a good starting point for a bootstrap understanding of an arbitrary linear transformation. First of all, if A is injective than V_0 = \{0_V\}, and the statement A(0_V)=0_W is just part of what it means to be a linear transformation. Furthermore, even in the case where V_0=\mathrm{Ker}(A) is nontrivial we learn nothing, and attempting to bootstrap from V_0 to all of V just brings us back to the injective situation, with no drop in complexity. The image of V_0 under A is the zero space W_0=\{0_W\} in W. Decomposing

V = V_0 \oplus V_0^\perp \quad\text{and}\quad W=W_0 \oplus W_0^\perp,

we tautologically have that A maps V_0^\perp into W_0^\perp=W. Restricting A to a map A \in \mathrm{Hom}(V_0^\perp,W), we are back in the injective situation. The dimension of the codomain W did not go down, and neither did the rank of A. In terms of matrices, if we choose an orthonormal basis X \subset V_0 of latex $V_0 = \mathrm{Ker}(A)$ and an orthonormal basis X^\perp \subset V_0^\perp, and take Y \subset W to be any orthonormal basis, then with respect to the bases X \sqcup X^\perp and Y we have

[A] =\begin{bmatrix} 0 & [A]\end{bmatrix},

where the row echelon form of the nested [A] has the same row echelon form as the external [A] up to the removal of zero columns.

The basic lesson to be learned from this failure is that in order to implement our bootstrap strategy we need to identify an exactly solvable subspace V_1 canonically associated to every morphism \mathrm{Hom}(V,W) by means to algebra and geometry, not just pure algebra. One such space is V_1=\mathrm{Opt}(A), the space

V_1 = \{v \in V \colon \|Av\|=\sigma_1\|v\|\}

whose nonzero vectors maximize the norm ratio \frac{\|Av\|}{\|v||}, a quantity bounded above by the operator norm \sigma_1=\|A\|. At the beginning of this week’s lectures we proved the following nontrivial theorem using the Parallelogram Law.

Proposition 16.1. The set V_1 is a nonzero subspace of V.

The key idea of the Singular Value Decomposition is that V_1 is a space on which the behavior of A can be completely understood, and this understanding can be bootstrapped to all of V. That is really what I want you to appreciate at a level beyond formulas. An arbitrary linear transformation can be completely understood in terms of the vectors which it maximally stretches.

First, let us explain why A is “exactly solvable” on V_1, no matter what A is. There are really two distinct cases: \sigma_1=0 and \sigma_1>0. The first case means A \in \mathrm{Hom}(V,W) is the zero transformation, and nothing more needs to be said. In the second case, \sigma_1>0, let W_1 be the image of V_1 under A and consider the restriction of A to V_1, which is a morphism A_1 \in \mathrm{Hom}(V_1,W_1). Since W_1=A(V_1), the map A_1 \in \mathrm{Hom}(V_1,W_1) is surjective by definition, and furthermore \|A_1\|=\sigma_1. Let us write A_1 \in \mathrm{Hom}(V_1,W_1) in the form A=\sigma_1 U_1 where U_1=\frac{1}{\sigma_1}A. Since U_1 is a nonzero scalar multiple of A_1, it is also surjective. Moreover, U_1 has operator norm equal to one, so it is an isometry and hence injective. Thus, U_1 \in \mathrm{Hom}(V_1,W_1) is an isometric isomorphism: the image of any orthonormal basis of X_1 \subset V_1 under U_1 is an orthonormal basis of Y_1 \subset W_1, and the matrix of U_1 \in \mathrm{Hom}(V_1,W_1) with respect to these bases is the square identity matrix

[U_1] = \begin{bmatrix} 1 & {} & {} \\ {} & \ddots & {} \\ {} & {} & 1\end{bmatrix}.

Thus, the morphism A_1 \in \mathrm{Hom}(V_1,W_1) with respect to the bases X_1 \subset V_1 and Y_1 \subset W_1 is

[A_1] = \begin{bmatrix} 1 & {} & {} \\ {} & \ddots & {} \\ {} & {} & 1\end{bmatrix},

so considering the restriction of A to V_1=\mathrm{Opt}(A) to be “exactly solved” should be noncontroversial. If you want a Euclidean image to hold in your mind, one option is the following: you can imagine V_1 and W_1 as two copies of the same space, and visualize A_1=\sigma_1 U_1 as mapping V_1 onto W_1 by first performing a rotation and then a dilation.

Now we want to bootstrap our local understanding of A \neq 0 on V_1 to a global understanding of A on V. It may happen that we get lucky, and no bootstrapping is required: V_1=V. This means exactly that the transformation A \in \mathrm{Hom}(V,W) is of the form A=\sigma U with \sigma>0 and U \in \mathrm{Hom}(V,W) an isometric isomorphism.

We now proceed with the case where V_1=\mathrm{Opt}(A) is a proper subspace of V (note that this implies A \in \mathrm{Hom}(V,W) is not the zero transformation). In this case we have to figure out how A acts on the complementary space V_1^\perp, which is not the zero subspace of V. Structurally, this means we need to figure out how A behaves with respect to the decompositions

V=V_1 \oplus V_1^\perp \quad\text{and}\quad W=W_1 \oplus W_1^\perp,

where by definition W_1=A(V_1) is the image of V_1=\mathrm{Opt}(A) under A. A good starting point is to wonder how V_0=\mathrm{Ker}(A) relates to the decomposition V=V_1 \oplus V_1^\perp, which is a point of intrinsic interest: what is the relationship between the vectors which A maximally stretches and those which it annihilates? It is easy to see that V_0 \cap V_1=\{0_V\}. Indeed, if v \in V_0 \cap V_1 then we have

Av=0_W \quad\text{and}\quad \|Av\|=\sigma_1\|v\|,

and since \sigma_1>0 the condition \sigma_1\|v\|=0 forces \|v\|=0.

Proposition 16.2. We have V_0 \leq V_1^\perp.

Proof: A more geometrically suggestive equivalent statement is to take the complement and write the claim as V_1 \leq V_0^\perp. This says that any vector in V which is maximally stretched by A is orthogonal to every vector annihilated by A, and indeed this can be established using elementary geometric reasoning. (Note to self: somehow I never noticed this striking fact before – this is a good exam problem). \square

The counterpart of Proposition 16.3 in the target space decomposition W=W_1 \oplus W_1^\perp is very straightforward: since W_1 = A(V_1) we have W_1 \leq \mathrm{Im}(A) which is equivalent to \mathrm{Im}(A)^\perp \leq W_1^\perp.

Now consider the extreme case of Proposition 16.2 where V_0=\mathrm{Ker}(A) exhausts V_1^\perp = \mathrm{Opt}(A)^\perp. Then, our decompositions

V= V_1 \oplus V_1^\perp \quad\text{and}\quad W=W_1 \oplus W_1^\perp

becomes

V=\mathrm{Opt}(A) \oplus \mathrm{Ker}(A) \quad\text{and}\quad W=\mathrm{Im}(A) \oplus \mathrm{Im}(A)^\perp,

and we have achieved our goal of understanding how A \in \mathrm{Hom}(V,W) acts on all of its domain V – it acts as the scaled isometric isomorphism \sigma_1U_1 in V_1=\mathrm{Opt}(A), and it acts as the zero operator in V_1^\perp = \mathrm{Ker}(A). This situation is important enough to warrant special terminology.

Definition 16.3. A morphism U \in \mathrm{Hom}(V,W) is said to be a partial isometry if it acts isometrically on a nonzero subspace V_1 of V and annihilates the complement V_1^\perp.

Equivalently, U \in \mathrm{Hom}(V,W) is a partial isometry if \|U\|=1 and

V=\mathrm{Opt}(U) \oplus \mathrm{Ker}(U).

What we have shown is that any nonzero morphism A \in \mathrm{Hom}(V,W) such that \mathrm{Opt}(A) and \mathrm{Ker}(A) are complementary subspaces of V is a positive multiple of a partial isometry.

The extremal case of Proposition 16.2 is just that – an extreme case. In general V_0=\mathrm{Ker}(A) may be a proper subspace of V_1^\perp = \mathrm{Opt}(A)^\perp and it remains to analyze this situation. This is where the following result (which we already proved using a perturbation argument) is essential.

Proposition 16.4. We have A(V_1)^\perp \leq W_1^\perp.

The force of proposition 16.4 is that in conjunction with the decompositions

V=V_1 \oplus V_1^\perp \quad\text{and}\quad W = W_1 \oplus W_1^\perp

it reduces our problem to a genuinely lower-complexity version of the same problem: we have

A = A_1 \oplus A_1^\perp,

where A_1 \in \mathrm{Hom}(V_1,W_1) is a morphism of the form A_1=\sigma_1U_1 with \sigma_1>0 and U_1 \in \mathrm{Hom}(V_1,W_1) an isometric isomorphism A_1^\perp \in \mathrm{Hom}(V_1^\perp,W_1^\perp) which remains to be analyzed, but has lower complexity than our original morphism A \in \mathrm{Hom}(V,W) because

\mathrm{rank}(A) = \mathrm{rank}(A_1) + \mathrm{rank}(A_1^\perp) \implies \mathrm{rank}(A_1^\perp) = \mathrm{rank}(A) - \dim V_1.

We now repeat our analysis of A \in \mathrm{Hom}(V,W) for the lower-complexity situation A_1^\perp \in \mathrm{Hom}(V_1^\perp,W_1^\perp), where both the dimension of the domain and the rank of the morphism have strictly decreased – this is what failed to happen when we tried to bootstrap starting with V_0=\mathrm{Ker}(A).

Let us carry out this iteration. Let V_2 = \mathrm{Opt}(A_1^\perp), i.e.

V_2 = \{v \in V_1^\perp \colon \|A_1^\perp v\|=\sigma_2\|v\|\}

where \sigma_2 = \|A_1^\perp\|. Note that since A_1^\perp \in \mathrm{Hom}(V_1^\perp,W_1^\perp) is nothing but the restriction of A \in \mathrm{Hom}(V,W) to V_1^\perp = \mathrm{Opt}(A)^\perp, we have \sigma_2 < \sigma_1 and the inequality is strict, as pointed out to me by Evan Young at least twice. If \sigma_2=0 we are done: A_1^\perp is the zero operator and we are back in the situation where A \in \mathrm{Hom}(V,W) is a positive multiple of a partial isometry U \in \mathrm{Hom}(V,W). Otherwise, \sigma_2 >0 and we define W_2 = A(V_2). Then, the restriction of A_1^\perp to W_2 is a morphism A_2 \in \mathrm{Hom}(V_2,W_2) of the form A_2=\sigma_2 U_2 with U_2 \in \mathrm{Hom}(V_2,W_2) an isometric isomorphism. As with our first pass, it remains to understand how A_1^\perp acts on the orthogonal complement of V_2 in V_1^\perp. Structurally, we need to understand the decompositions

V_1^\perp = V_2 \oplus V_2^\perp \quad\text{and}\quad W_1^\perp = W_2 \oplus W_2^\perp.

As in our first iteration, \mathrm{Ker}(A_1^\perp) is a subspace of V_2^\perp=\mathrm{Opt}(A_1^\perp)^\perp. If \mathrm{Ker}(A_1^\perp) exhausts V_2^\perp then we are done: we have

V_1^\perp = V_2 \oplus \mathrm{Ker}(A_1^\perp) \quad\text{and}\quad W_1^\perp = W_2 \oplus \mathrm{Im}(A_1^\perp)^\perp,

which lifts back up to the level of A \in \mathrm{Hom}(V,W) as

V= V_1 \oplus V_2 \oplus \mathrm{Ker}(A) \quad\text{and}\quad W=W_1 \oplus W_1 \oplus \mathrm{Im}(A)^\perp.

If not, we repeat the process on the restriction of A_1^\perp to V_2^\perp, which is a morphism A_3 \in \mathrm{Hom}(V_2^\perp,W_2^\perp) of rank

\mathrm{rank}(A_3)=\mathrm{rank}(A)-\dim V_1 - \dim V_2.

Continuing the process until we exhaust \mathrm{rank}(A) \leq \min(\dim V,\dim W), we arrive at the Singular Value Decomposition.

Theorem 16.5. (SVD, chunky geometric form) For any morphism A \in \mathrm{Hom}(V,W), we have orthogonal decompositions

V = V_1 \oplus \dots V_k \oplus \mathrm{Ker}(A) \quad\text{and}\quad W=W_1 \oplus \dots \oplus W_k \oplus \mathrm{Im}(A)^\perp

and numbers

\sigma_1 > \dots > \sigma_k >0

such that, for each 1 \leq i \leq k, the spaces V_i,W_i are nonzero and the restriction of A to V_i is a morphism A_i \in \mathrm{Hom}(V_i,W_i) of the form A_i=\sigma_iU_i with U_i \in \mathrm{Hom}(V_i,W_i) an isometric isomorphism.

The paired spaces V_1,W_1,\dots,V_k,W_k in the SVD are called the left and right singular spaces of A. Although these spaces are nonzero, the number k of singular spaces may be zero. Indeed, \dim(V_1)+\dots+\dim(V_k)=\mathrm{rank}(A), and the case k=0 occurs precisely when A \in \mathrm{Hom}(V,W) is the zero transformation.

In general, the worst case scenario for the runtime of the recursive routine which establishes the SVD occurs when A \in \mathrm{Hom}(V,W) has a one-dimensional optimizer space V_1=\mathrm{Opt}(A), and the restriction of A to V_1^\perp again has a one-dimensional optimizer space, and so on until the rank r of A is exhausted after r iterations. In this case, A has r pairs of one-dimensional singular spaces V_1,W_1,\dots,V_r,W_r, and r singular values \sigma_1 > \dots > \sigma_r>0.

Even when the worst case scenario is not what occurs, we can write the Singular Value Decomposition in a form where all subspaces involved are one-dimensional, simply by choosing a basis within each V_i and writing it as a direct sum of lines. This leads to an alternative formulation of the SVD with possibly repeated singular values, which really means singular values with multiplicity.

Theorem 16.6. (SVD, fine-grained geometric version) For any morphism A \in \mathrm{Hom}(V,W) of rank r, we have orthogonal decompositions

V = V_1 \oplus \dots V_r \oplus \mathrm{Ker}(A) \quad\text{and}\quad W=W_1 \oplus \dots \oplus W_r \oplus \mathrm{Im}(A)^\perp

and numbers

\sigma_1 \geq \dots \geq \sigma_r >0

such that, for each 1 \leq i \leq r, the spaces V_i,W_i are one-dimensional and the restriction of A to V_i is a morphism A_i \in \mathrm{Hom}(V_i,W_i) of the form A_i=\sigma_iU_i with U_i \in \mathrm{Hom}(V_i,W_i) an isometric isomorphism of the lines V_i and W_i.

What is confusing in the way the two versions of the SVD are stated is that the symbols V_i,W_i represent different things, but these things are not unrelated. Nevertheless I will leave our two versions as stated above instead of introducing different letters for each version.

Leave a Reply