ISBN: 0-387-95318-3

Текст
                    Matrix Algebra:
Exercises and
Solutions
David A. Harville
Springer


Matrix Algebra: Exercises and Solutions
Springer New York Berlin Heidelberg Barcelona Hong Kong London Milan Paris Singapore Tokyo
David A. Harville Matrix Algebra: Exercises and Solutions Springer
David A. Harville Mathematical Sciences Department IBM T.J. Watson Research Center Yorktown Heights, NY 10598-0218 USA Library of Congress Cataloging-in-Publication Data Harville, David A. Matrix algebra: exercises and solutions / David A. Harville. p. cm. Includes bibliographical references and index. ISBN 0-387-95318-3 (pbk.: alk. paper) 1. Matrices—Problems, exercises, etc. I. Title. QA188.H38 2001 519.9*434—dc21 2001032838 Printed on acid-free paper. © 2001 Springer-Verlag New York, Inc. All rights reserved. This work may not be translated or copied in whole or in part without the written permission of the publisher (Springer-Verlag New York, Inc., 175 Fifth Avenue. New York, NY 10010, USA), except for brief excerpts in connection with reviews or scholarly analysis. Use in connection with any form of information storage and retrieval, electronic adaptation, computer software, or by similar or dissimilar methodology now known or hereafter developed is forbidden. The use of general descriptive names, trade names, trademarks, etc., in this publication, even if the former are not especially identified, is not to be taken as a sign that such names, as understood by the Trade Marks and Merchandise Marks Act, may accordingly be used freely by anyone. Production managed by Yong-Soon Hwang; manufacturing supervised by Jeffrey Taub. Photocomposed copy prepared from the author's LaTeX file. Printed and bound by Maple-Vail Book Manufacturing Group, York, PA. Printed in the United States of America. 987654321 ISBN 0-387-95318-3 SPIN 10841733 Springer-Verlag New York Berlin Heidelberg A member of BertelsmannSpringer Science*Business Media GmbH
Preface This book comprises well over three-hundred exercises in matrix algebra and their solutions. The exercises are taken from my earlier book Matrix Algebra From a Statistician's Perspective. They have been restated (as necessary) to make them comprehensible independently of their source. To further insure that the restated exercises have this stand-alone property, I have included in the front matter a section on terminology and another on notation. These sections provide definitions, descriptions, comments, or explanatory material pertaining to certain terms and notational symbols and conventions from Matrix Algebra From a Statistician's Perspective that may be unfamiliar to a nonreader of that book or that may differ in generality or other respects from those to which he/she is accustomed. For example, the section on terminology includes an entry for scalar and one for matrix. These are standard terms, but their use herein (and in Matrix Algebra From a Statistician's Perspective) is restricted to real numbers and to rectangular arrays of real numbers, whereas in various other presentations, a scalar may be a complex number or more generally a member of a field, and a matrix may be a rectangular array of such entities. It is my intention that Matrix Algebra: Exercises and Solutions serve not only as a "solution manual" for the readers of Matrix Algebra From a Statistician's Perspective, but also as a resource for anyone with an interest in matrix algebra (including teachers and students of the subject) who may have a need for exercises accompanied by solutions. The early chapters of this volume contain a relatively small number of exercises—in fact. Chapter 7 contains only one exercise and Chapter 3 only two. This is because the corresponding chapters of Matrix Algebra From a Statistician s Perspective cover relatively standard material, to which many readers will have had previous exposure, and/or are relatively short. It is
vi Preface the final ten chapters that contain the vast majority of the exercises. The topics of many of these chapters are ones that may not be covered extensively (if at all) in more standard presentations or that may be covered from a different perspective. Consequently, the overlap between the exercises from Matrix Algebra From a Statistician's Perspective (and contained herein) and those available from other sources is relatively small. A considerable number of the exercises consist of verifying or deriving results supplementary to those included in the primary coverage of Matrix Algebra From a Statistician's Perspective. Thus, their solutions provide what are in effect proofs. For many of these results, including some of considerable relevance and interest in statistics and related disciplines, proofs have heretofore only been available (if at all) through relatively high-level books or through journal articles. The exercises are arranged in 22 chapters and within each chapter, are numbered successively (starting with 1). The arrangement, the numbering, and the chapter titles match those in Matrix Algebra From a Statistician's Perspective. An exercise from a different chapter is identified by a number obtained by inserting the chapter number (and a decimal point) in front of the exercise number. A considerable effort was expended in designing the exercises to insure an appropriate level of difficulty—the book Matrix Algebra From a Statistician's Perspective is essentially a self-contained treatise on matrix algebra, however it is aimed at a reader who has had at least some previous exposure to the subject (of the kind that might be attained in an introductory course on matrix or linear algebra). This effort included breaking some of the more difficult exercises into relatively palatable parts and/or providing judicious hints. The solutions presented herein are ones that should be comprehensible to those with exposure to the material presented in the corresponding chapter of Matrix Algebra From a Statistician's Perspective (and possibly to that presented in one or more earlier chapters). When deemed helpful in comprehending a solution, references are included to the appropriate results in Matrix Algebra From a Statistician 's Perspective—unless otherwise indicated a reference to a chapter, section, or subsection or to a numbered result (theorem, lemma, corollary, "equation", etc.) pertains to a chapter, section, or subsection or to a numbered result in Matrix- Algebra From a Statistician's Perspective (and is made by following the same conventions as in the corresponding chapter of Matrix Algebra From a Statistician's Perspective). What constitutes a "legitimate" solution to an exercise depends of course on what one takes to be "given". If additional results are regarded as given, then additional, possibly shorter solutions may become possible. The ordering of topics in Matrix Algebra From a Statistician's Perspective is somewhat nonstandard. In particular, the topic of eigenvalues and eigenvectors is deferred until Chapter 21, which is the next-to-last chapter. Among the key results on that topic is the existence of something called the spectral decomposition. This result if included among those regarded as given, could be used to devise alternative solutions for a number of the exercises in the chapters preceding Chapter 21. However, its use comes at a "price"; the existence of the spectral decomposition can only be established by resort to mathematics considerably deeper than those
Preface vii underlying the results of Chapters 1-20 in Matrix Algebra From a Statistician's Perspective. I am indebted to Emmanuel Yashchin for his support and encouragement in the preparation of the manuscript for Matrix Algebra: Exercises and Solutions. I am also indebted to Lorraine Menna, who entered much of the manuscript in I#IjeX, and to Barbara White, who participated in the latter stages of the entry process. Finally, I wish to thank John Kimmel, who has been my editor at Springer-Verlag, for his help and advice.
Contents Preface v Some Notation xi Some Terminology xvii 1 Matrices 1 2 Submatrices and Partitioned Matrices 7 3 Linear Dependence and Independence 11 4 Linear Spaces: Row and Column Spaces 13 5 Trace of a (Square) Matrix 19 6 Geometrical Considerations 21 7 Linear Systems: Consistency and Compatibility 27 8 Inverse Matrices 29 9 Generalized Inverses 35 10 Idempotent Matrices 49
x Contents 11 Linear Systems: Solutions 55 12 Projections and Projection Matrices 63 13 Determinants 69 14 Linear, Bilinear, and Quadratic Forms 79 15 Matrix Differentiation 113 16 Kronecker Products and the Vec and Vech Operators 139 17 Intersections and Sums of Subspaces 161 18 Sums (and Differences) of Matrices 179 19 Minimization of a Second-Degree Polynomial (in n Variables) Subject to Linear Constraints 209 20 The Moore-Penrose Inverse 221 21 Eigenvalues and Eigenvectors 231 22 Linear Transformations 251 References 265 Index 267
Some Notation {*;} A row or (depending on the context) column vector whose ith element isxt [aij) A matrix whose //th element is a\} (and whose dimensions are arbitrary or may be inferred from the context) A' The transpose of a matrix A Ap The /?th (for a positive integer p) power of a square matrix A; i.e., the matrix product AA • • ■ A defined recursively by setting A0 = I and taking A* = AA*"1 (* = 1 p) C(A) Column space of a matrix A 71( A) Row space of a matrix A TZmxn The linear space comprising all m x n matrices TV The linear space TZ"xl comprising all w-dimensional column vectors or (depending on the context) the linear space 7£lx,! comprising all /i- dimensional row vectors sp(S) Span of a finite set S of matrices; sp({Ai,..., A*}), which represents the span of the set {Ai, , Ajt} comprising the k matrices Ai,..., Ajt, is generally abbreviated to sp(Ai,..., A*) C Writing S C T (or T D S) indicates that a set S is a (not necessarily proper) subset of a set T dim(V) Dimension of a linear space V rank A The rank of a matrix A rank T The rank of a linear transformation T
Xll Some Notation tr(A) The trace of a (square) matrix A 0 The scalar zero (or, depending on the context, the zero transformation from one linear space into another) 0 A null matrix (whose dimensions are arbitrary or may be inferred from the context) / An identity transformation 1 An identity matrix (whose order is arbitrary or may be inferred from the context) I„ An identity matrix of order n A • B Inner product of a pair of matrices A and B (or if so indicated, quasi-inner product of the pair A and B) || A || Norm of a matrix A (or, in the case of a quasi-inner product, the quasi norm of A) 5(A, B) Distance between two matrices A and B T~l The inverse of an invertible transformation T A~l The inverse of an invertible matrix A A~ An arbitrary generalized inverse of a matrix A .AAA) Null space of a matrix A N( T) Null space of a linear transformation T _L A symbol for "is orthogonal to" _L\v A symbol used to indicate (by writing x ±w y. x ±vv U, or U JLw V) that 2 vectors x and y, a vector x and a subspace U* or 2 subspaces U and V are orthogonal with respect to a symmetric nonnegative definite matrix W Px The matrix X(X'X)~X' [which is invariant to the choice of the generalized inverse (X'X)~] Px.w The matrix X(X'WX)~X'W [which if W is symmetric and positive definite, is invariant to the choice of the generalized inverse (X'WX)-] U1 The orthogonal complement of a subspace U of a linear space V CL(X) The orthogonal (with respect to the usual inner product) complement of the column space C(X) of an n x p matrix X [when C{X) is regarded as a subspace of 7^"] CW(X) The orthogonal complement of the column space C(X) of an nxp matrix X when, for an n x n symmetric positive definite matrix W, the inner product is taken to be the bilinear form x'Wy [and when C{X) is regarded as a subspace of VJ'] cr„{*) A function whose value <7„(/|, j\\ ...;/„, j„) for any two (not necessarily different) permutations of the first n positive integers is the number of negative pairs among the (") pairs that can be formed from the i\j\ /ff./fith elements of an n x n matrix 0„(«) A function whose value 0„(/| /„) for any sequence of n distinct
Some Notation xiii integers i\ i„ is p\ -\ h pn-u where (fork = 1,...,/1-1) /¾ represents the number of integers in the subsequence /*+i i„ that are smaller than /* | A | The determinant of a square matrix A — with regard to partitioned matrices, A,,.\ An ... Au. An \ \ \ 11 may be abbreviated to | \Arl ... Arc) det( A) The determinant of a square matrix A adj(A) The adjoint matrix of a square matrix A J A matrix, all of whose elements equal one (and whose dimensions are arbitrary or may be inferred from the context) J,»„ An 777 x 77 matrix, all of whose 77777 elements equal one A <g> B The Kronecker product of a matrix A and a matrix B — this notation extends in an obvious way to the Kronecker product of 3 or more matrices vec A The vec of a matrix A vech A The vech of a (square) matrix A K„m The mn x 77777 vec-permutation (or commutation) matrix G„ The 772 x 77(77 + 1)/2 duplication matrix H„ An arbitrary left inverse of G„, so that H„ is any /7(/7 + 1)/2 x n2 matrix such that H„G„ = I or equivalently such that, for every 77 x n symmetric matrix A, vech A = H„ vec A — one choice for H„ is H„ = (Gj,G„)~ Gj, Djf(c) The 7th (first-order) partial derivative of a function /, with domain S in Hmx *, at an interior point c of S — the function whose value at a point c is Djf(c) is represented by the symbol Djf tt4t The jth partial derivative of a function / of an m x 1 vector x = Ui,..., a,,,)' — an alternative notation to Djf or Djf(x) D/(c) The 1 x 7/7 vector [£>i/(c) An(0] (where / is a function with domain S in 7£mxl and where c is an interior point of S)— similarly, D/ is the 1 x 177 vector (D\f%..., Dmf) ^- The 777 x 1 vector (df/Bx 1 Bf/dx„,)' of partial derivatives of a function / of an /77 x 1 vector x = (.r 1 .vlw)' — an alternative [to (D/)' or (D/(x))'] notation for the gradient vector |4 The 1 x //7 vector (Bf/dxi df/dx„,) of partial derivatives of a function / of an 777 x 1 vector x = (x\ xm)' — equals (df/dxY and is an alternative notation to D/ or D/(x) D2./(c) The 77 th second-order partial derivative of a function /, with domain S in ft'" x l, at an interior point c of S—the function whose value at a point c is Dfjf(c) is represented by the symbol D?./
xiv Some Notation d2f fix.Qx. An alternative [to D?. / or Z& /(x)] notation for the 17th (second-order) partial derivative of a function f of an m x 1 vector x = (x\,..., xm)'— this notation extends in a straightforward way to third- and higher-order partial derivatives H/ The Hessian matrix of a function / — accordingly, H/(c) represents the value of H/ at an interior point c of the domain of / Djf The p x 1 vector (Dj f\,..., Djfp)', whose ith element is the 7 th partial derivative Dj ft of the 1 th element fi of a p x 1 vector f = (/1,..., fpY of functions, each of whose domain is a set S in 1lmxl — similarly, Djf(c) = [Djfi (c),..., Djfp(c)Y, where c is an interior point of S ^7 The p x q matrix whose sf th element is the partial derivative dfst/dxj of the stih element of a p x q matrix F = [fst] of functions of a vector x = (x\,..., xmY of m variables fa., ay. The P x# matrix whose .sf th element is the second-order partial derivative d2fst/d*idxj of the stth element of a p xq matrix F = {/^} of functions of a vector x = (*i,..., xm)' of m variables—this notation extends in a straightforward way to a p x q matrix whose stth element is one of the third- or higher-order partial derivatives of the sf th element of F Df The Jacobian matrix (Dif,..., Dmt) of a vector f = (/1 fp)' of functions, each of whose domain is a set S in 7£mxl — similarly, Df(c) = [£>tf(c) Djnf(c)], where c is an interior point of S 7¾ An alternative [to Df orDf(x)] notation for the Jacobian matrix of a vector f = (/1 fp)' of functions of an m x 1 vector x = (*i xm)' — df/dtf is the p x m matrix whose ij\h element is dfi/dxj ^ An alternative [to (Df/ or (Df(x)/] notation for the gradient (matrix) of a vector f = (/1 fp)' of functions of an m x 1 vector x = (jt|,...,jc„,)'— df'/dx is the m x p matrix whose 71th element is Bfi/dxj ■*h The derivative of a function / of an m x n matrix Xoimn "independent" variables or (depending on the context) of an n x n symmetric matrix X — the matrix df/dX' is identical to (df/dX)' UC\V The intersection of 2 sets U and V of matrices—this notation extends in an obvious way to the intersection of 3 or more sets UUV The union of 2 sets U and V of matrices (of the same dimensions)—this notation extends in an obvious way to the union of 3 or more sets U + V The sum of 2 nonempty sets U and V of matrices (of the same dimensions)—this notation extends in an obvious way to the sum of 3 or more nonempty sets U © V The direct sum of 2 (essentially disjoint) linear spaces U and V in %m xn
Some Notation xv — writing U © V (rather than U + V) serves to emphasize, or (in the absence of any previous indication) imply, that U and V are essentially disjoint and hence that their sum is a direct sum A+ The Moore-Penrose inverse of a matrix A (kT) The scalar multiple of a scalar k and a transformation T from a linear space V into a linear space W; in the absence of any ambiguity, the parentheses may be dropped, that is, kT may be written in place of (kT) (T + S) The sum of two transformations T and S from a linear space V into a linear space W; in the absence of any ambiguity, the parentheses may be dropped, that is, T + S may be written in place of (T+S) — this notation extends in an obvious way to the sum of three or more transformations (TS) The product of a transformation T from a linear space V into a linear space W and a transformation S from a linear space U into V; in the absence of any ambiguity, the parentheses may be dropped, that is, TS may be written in place of (TS) — this notation extends in an obvious way to the product of three or more transformations Lb A transformation defined for any (nonempty) linearly independent set B of matrices (of the same dimensions), say the matrices Yi, Y2 Y„: it is the transformation from 7£"xl onto the linear space W = sp(#) that assigns to each vector x = (xi,*2» • • • .*«)' in ^nxl the matrix jciYi + x2Y2 + • • • + xnYn in VV.
Some Terminology adjoint matrix The adjoint matrix of an n x n matrix A = [an} is the transpose of the cofactor matrix of A (or equivalently is the n x n matrix whose ijth element is the cofactor ofay,). algebraic multiplicity The characteristic polynomial, say p(»), of an n x n matrix A has a unique (aside from the order of the factors) representation of the form p{\) = (-!)"(* - Xi)" ■ • • (X - Xk)Ykq(X) (-00 < X < oo), where {Xj X*} is the spectrum of A (comprising the distinct scalars that are eigenvalues of A), y\ yjt are (strictly) positive integers, and q is a polynomial (of degree n - Y^a=\ Yi)tnat nas n0 real roots; for / = I A:, Yi is referred to as the algebraic multiplicity of the eigenvalue X,-. basis A basis for a linear space V is a finite set of linearly independent matrices in V that spans V. basis (natural) The natural basis for Hmx" comprises the mn matrices Un, U21, .. .,Umi,..., Ui„,U2n Umw, where (for/ = I,..., m andy = 1,... n) Vij is the m x n matrix whose ijth element equals 1 and whose remaining mn — 1 elements equal 0; the natural (or usual) basis for the linear space of all n x n symmetric matrices comprises the n(n + 1)/2 matrices U^, U^, • • •. U*, 1¾. «;+, j Ki Vm, where (for , = 1....,"») 1¾ is the n x n matrix whose ith diagonal element equals 1 and whose remaining n2 - 1 elements equal 0 and (for j < i = I n) \J*j is the /7 x n matrix whose ijth and jith elements equal 1 and whose remaining n2 — 2 elements equal 0. bilinear form A bilinear form in an m x 1 vector x = (.vi .*„,)' and an n x 1 vector y = (y\ v,,)' is a function of x and y (defined for x € fcm and
Some Terminology y € ft") that, for some mxn matrix A = [ay) (called the matrix of the bilinear form), is expressible as x'Ay = £,. : a^xiyj — the bilinear form is said to be symmetric if m = n and x'Ay = /Ax for all x and all y or equivalently if the matrix A is symmetric. /An 0 ... 0\ block-diagonal A partitioned matrix of the form 0 A22 (all \ 0 0 \rr) of whose off-diagonal blocks are null matrices) is said to be block-diagonal and may be expressed in abbreviated notation as diag(An, A22 Arr). /An A12 ... AiA block-triangular A partitioned matrix of the form 0 A22 /An 0 A21 A22 0\ 0 A2r Arr/ is respectively upper or lower block-triangular- \Ari Ar2 Arr/ to indicate that a partitioned matrix is upper or lower block-triangular (without specifying which), the matrix is referred to simply as block-triangular. characteristic polynomial (and equation) Corresponding to any n x n matrix A is its characteristic polynomial, say /?(•), defined (for -00 < k < 00) by p(X) = IA — XI|, and its characteristic equation p(X) = 0 obtained by setting its characteristic polynomial equal to 0; p(X) is a polynomial in X of degreen and hence is of the form p{X) = cq+ciX-\ \rCn-iX"~l+cnX", where the coefficients <?o, ci,..., c„_i, c„ depend on the elements of A. Cholesky decomposition The Cholesky decomposition of a symmetric positive definite matrix, say A, is the unique decomposition of the form A = T'T, where T is an upper triangular matrix with positive diagonal elements. More generally, the Cholesky decomposition of an n xn symmetric nonnegative definite matrix, say A, of rank r is the unique decomposition of the form A = T'T, where T is an n x n upper triangular matrix with r positive diagonal elements and n — r null rows. cofactor (and minor) The cofactor and minor of the iyth element, say ay, of an nxn matrix A are defined in terms of the (n — 1) x (n - 1) submatrix, say A/y, of A obtained by striking out the /th row and jth column (i.e., the row and column containing ay): the minor of a,j is |A/y |, and the cofactor is the "signed" minor (-l),+y|Aly|. cofactor matrix The cofactor matrix (or matrix of cofactors) of an n x n matrix A = {aij} is the nxn matrix whose //th element is the cofactor of a-,}. column space The column space of an m x n matrix A is the set whose elements consist of all m-dimensional column vectors that are expressible as linear
Some Terminology xix combinations of the n columns of A. commute Two n x n matrices A and B are said to commute if AB = BA. commute in pairs n x n matrices, say Ai Ajt, are said to commute in pairs if A5A/ = A/A5 for s > i = 1,..., k. consistent A linear system is said to be consistent if it has one or more solutions. continuous A function /, with domain S in ft1"*1, is continuous at an interior point c of S if limx_c /(x) = /(c). continuously differentiable A function / , with domain S in Hmx], is continuously differentiable at an interior point c of S if £>i/(c), D2/(c), ..., Dmf{c) exist and are continuous at every point x in some neighborhood of c — a vector or matrix of functions is continuously differentiable at c if all of its elements are continuously differentiable at c. derivative of a function of a matrix The derivative of a function / of an m x n matrix X = [x;j) of mn "independent" variables is the/n x n matrix whose ijth element is the partial derivative df/dxij of / with respect to Xjj when / is regarded as a function of an m/i-dimensional column vector x formed from X by rearranging its elements; the derivative of a function / of an n x n symmetric (but otherwise unrestricted) matrix of variables is the n x n (symmetric) matrix whose ijth element is the partial derivative df/dxij or Bf/dxji of / with respect to x^ or jcy,- when / is regarded as a function of an n(n + l)/2-dimensional column vector x formed from any set of n(n +1 )/2 nonredundant elements of X. determinant The determinant of an n x n matrix A = {ay) is (by definition) the (scalar-valued) quantity £ (—1)*iO*i J^a\jt • • • a„jn, or equivalently the quantity £ (-l)<M'i '«>«,•, j... ainn, where j\ ;„ or /| i„ is a permutation of the first n positive integers and the summation is over all such permutations. diagonalization Annxn matrix, say A, is said to be diagonalizable if there exists an n x n nonsingular matrix Q such that Q"1 AQ is diagonal, in which case Q is said to diagonalize A (or A is said to be diagonalized by Q); a matrix that can be diagonalized by an orthogonal matrix is said to be orthogonally diagonalizable. diagonalization (simultaneous) k matrices, say Ai,..., A*, of dimensions nxn, are said to be simultaneously diagonalizable if all k of them can be diagonalized by the same matrix, that is, if there exists an n x n nonsingular matrix Q such that Q"1 AiQ,..., Q"1 A*Q are all diagonal, in which case Q is said to simultaneously diagonalize Aj A* (or Ai Ajt are said to be simultaneously diagonalized by Q). dimension (of a linear space) The dimension of a linear space V is the number of matrices in a basis for V. dimension (of a row or column vector) A row or column vector having n elements is said to be of dimension n.
XX Some Terminology dimensions (of a matrix) A matrix having m rows and n columns is said to be of dimensions m x n. direct sum If 2 linear spaces in Tlmx" are essentially disjoint, their sum is said to be a direct sum. distance The distance between two matrices A and B in a linear space V is II A-B ||. dual transformation Corresponding to any linear transformation T from an n- dimensional linear space V into an w»-dimensional linear space W is a linear transformation from W into V called the dual transformation: denoting by X • Z the inner product of an arbitrary pair of matrices X and Z in V and by U * Y the inner product of an arbitrary pair of matrices U and Y in W, the dual transformation is the (unique) linear transformation, say 5, from W into V such that (for every matrix X in V and every matrix Y in W) X-S(Y) = !T(X)*Y; further, for all Y in W,5(Y) = £'j=1 [Y*7,(Xy)]Xy, where Xj, X2 X„ are any matrices that form an orthonormal basis for V. duplication matrix The n2 x n(n +1 )/2 duplication matrix is the matrix, denoted by the symbol G„, such that, for every nxn symmetric matrix A, vec A = G„vech A. eigenspace The eigenspace of an eigenvalue, say A, of an n x n matrix A is the linear space jV(A — AI) — with the exception of the nx\ null vector, every member of this space is an eigenvector (of A) corresponding to A. eigenvalues and eigenvectors An eigenvalue of an n x n matrix A is (by definition) a scalar (real number), say A, for which there exists an n x 1 vector, say x, such that Ax = Ax, or equivalently such that (A — AI)x = 0; any such vector x is referred to as an eigenvector (of A) and is said to belong to (or correspond to) the eigenvalue A — eigenvalues (and eigenvectors), as defined herein, are restricted to real numbers (and vectors of real numbers). eigenvalues (not necessarily distinct) The characteristic polynomial, say /?(•), of an n x n matrix A is expressible as /7(A) = (-1)"(A - </i)(A - rf2) • • • (X - dm)q(X) (-00 < A < 00), where d\, di dm are not-necessarily-distinct scalars and q(») is a polynomial (of degree n —m) that has no real roots; d\, d: d,„ are referred to as the not-necessarily-distinct eigenvalues of A or (at the possible risk of confusion) simply as the eigenvalues of A—if the spectrum of A has k members, say A1 A*, with algebraic multiplicities of y\ y*, respectively, then /« = J^=\ Yh and (for i = 1 k) yi of the m not-necessarily-distinct eigenvalues equal A/. essentially disjoint Two subspaces, say U and V, of 7£'"x" are (by definition) essentially disjoint if U C\ V = {0}, i.e., if the only matrix they have in common is the (m x //) null matrix—note that every subspace of 1Z'"X" contains the (/11 x /0 null matrix, so that no two subspaces can be entirely disjoint.
Some Terminology xxi full column rank An m x n matrix A is said to have full column rank if rank(A) = 71. full row rank An m x n matrix A is said to have full row rank if rank(A) = in. generalized eigenvalue problem The generalized eigenvalue problem consists of finding, for a symmetric matrix A and a symmetric positive definite matrix B, the roots of the polynomial |A - XB| (i.e., the solutions for X to the equation |A - XB| = 0). generalized inverse A generalized inverse of an /» x /7 matrix A is any 72 x m matrix G such that AGA = A — if A is nonsingular, its only generalized inverse is A~l; otherwise, it has infinitely many generalized inverses. geometric multiplicity The geometric multiplicity of an eigenvalue, say X, of an /7 x n matrix A is (by definition) dim[Af(\ - XI)] (i.e., the dimension of the eigenspaceofX). gradient (or gradient matrix) The gradient of a vector f = (f\ fp)' of functions, each of whose domain is a set in 7£"'xl, is the m x p matrix [(D/i)' (DfpYl whose jith element is Djf;—the gradient of f is the transpose of the Jacobian matrix off. gradient vector The gradient vector of a function /, with domain in 7£"'xl, is the 772-dimensional column vector (D/)', whose jth element is the partial derivative Djf of / Hessian matrix The Hessian matrix of a function /, with domain in 1Zmxl, is the 777 x m matrix whose ijth element is the //th partial derivative Df.f of / homogeneous linear system A linear system (in a matrix X) of the form AX = 0; i.e., a linear system whose right side is a null matrix. idempotent A (square) matrix A is idempotent if A*" = A. identity transformation An identity transformation is a transformation from a linear space V onto V defined by T(X) = X. indefinite A square (symmetric or nonsymmetric) matrix or a quadratic form is (by definition) indefinite if it is neither nonnegative definite nor nonpositive definite—thus, an 77 x /7 matrix A and the quadratic form x'Ax (in an /? x 1 vector x) are indefinite if x'Ax < 0 for some x and x'Ax > 0 for some (other) x. inner product The inner product A »B of an arbitrary pair of matrices A and B in a linear space V is the value assigned to A and B by a designated function having the following 4 properties: (1) A -B = B • A; (2) A • A > 0, with equality holding if and only if A = 0; (3) (#VA) *B = /V(A-B) (where k is an arbitrary scalar); (4) (A + B) • C = (A • C) + (B • C) (where C is an arbitrary matrix in V)—the quasi-inner product A»B is defined in the same way as the inner product except that Property (2) is replaced by the weaker property (2') A* A > 0, with equality holding if A = 0. inner product (usual) The usual inner product of a pair of matrices A and B in a linear space is tr(A'B) (which in the special case of a pair of column vectors
xxii Some Terminology a and b reduces to a'b). interior point A matrix, say X, in a set S of m x n matrices is an interior point of S if there exists a neighborhood, say N, of X such that N C S. intersection The intersection of 2 sets, say U and V, of mxn matrices is the set comprising all matrices that are contained in both U and V; more generally, the intersection of k sets, say U\%..., £4, of m x n matrices is the set comprising all matrices that are contained in every one of U\ 24. invariant subspace A subspace U of the linear space 11" x l is said to be invariant relative to an n x n matrix A if, for every vector x in U, the vector Ax is also in U\ a subspace U of an ^-dimensional linear space V is said to be invariant relative to a linear transformation T from V into V if T(U) C U, that is, if the image T(U) of U is a subspace of U itself. inverse (matrix) A matrix B that is both a right and left inverse of a matrix A (so that AB = I and BA = I) is called an inverse of A. inverse (transformation) The inverse of an invertible transformation T from a linear space V into a linear space W is the transformation from W into V that assigns to each matrix Y in W the (unique) matrix X (in V) such that T(K) = Y. invertible (matrix) A matrix that has an inverse is said to be invertible—a matrix is invertible if and only if it is nonsingular. invertible (transformation) A transformation from a linear space V into a linear space W is (by definition) invertible if it is both 1-1 and onto. involutory A (square) matrix A is involutory if A2 = I, i.e., if it is invertible and is its own inverse. isomorphic If there exists a 1-1 linear transformation, say 7\ from a linear space V onto a linear space W, then V and W are said to be isomorphic, and T is said to be an isomorphism of V onto W. Jacobian matrix The Jacobian matrix of a p-dimensional vector f = (/i fpY of functions, each of whose domain is a set in 7£"'xl, is the p x m matrix (Dif,.... D,„f), whose ijth element is Djfi — in the special case where p = m, the determinant of this matrix is referred to as the Jacobian (or Jacobian determinant) off. Kronecker product The Kronecker product of two matrices, say an m x n matrix A = [ajj] and a p x q matrix B, is the mp x nq matrix /fluB #i2B ... «i„B\ I «21B «22B ... «2nB I \flmiB a„aK ... a,„„B/ obtained by replacing each element a-Xj of A with the p x q matrix tf,yB — the Kronecker-product operation is associative [for any 3 matrices A, B, and C, A <g> (B <g> C) = (A <g> B) <g> CI, so that the notion of a Kronecker product extends in an unambiguous way to 3 or more matrices.
Some Terminology XXIII k times continuously differentiate A function /, with domain S in ftmxl, is k times continuously differentiable at an interior point c of S if it and all of its first- through (k — l)th-order partial derivatives are continuously differentiable at c or, equivalently, if all of the first- through fcth-order partial derivatives of / exist and are continuous at every point in some neighborhood of c—a vector or matrix of functions is k times continuously differentiable at c if all of its elements are k times continuously differentiable at c. LDU decomposition An LDU decomposition of a square matrix, say A, is a decomposition of the form A = LDU, where L is a unit lower triangular matrix, D a diagonal matrix, and U an upper triangular matrix. least squares generalized inverse A generalized inverse, say G, of an m x n matrix A is said to be a least squares generalized inverse (of A) if (AG)' = AG; or, equivalently, an n x m matrix is a least squares generalized inverse of A if it satisfies Moore-Penrose Conditions (1) and (3). left inverse A left inverse of an m x n matrix A is an n x m matrix L such that LA = I„ — a matrix has a left inverse if and only if it has full column rank. linear dependence or independence A nonempty (but finite) set of matrices (of the same dimensions), say Ai, A2,..., Ajt, is (by definition) linearly dependent if there exist scalars x\, X2 **, not all 0, such that £f=1 *,-A/ = 0; otherwise (if no such scalars exist), the set is linearly independent—by convention, the empty set is linearly independent. linear space The use of this term is confined (herein) to sets of matrices (all of which have the same dimensions). A nonempty set, say V, is called a linear space if: (1) for every matrix A in V and every matrix B in V, the sum A + B is in V; and (2) for every matrix A in V and every scalar /:, the product kA is in V. linear system A linear system is (for some positive integers m,«, and p) a set of mp simultaneous equations expressible in nonmatrix form as £"=1 a\jXjk = bjk (i =/ m; k = 1 p), or in matrix form as AX = B, where A = [aij] is an m x n matrix comprising the "coefficients", X = [xjt] is an n x p matrix comprising the "unknowns", and B = (½} is an m x p matrix comprising the "right (hand) sides"—A is referred to as the coefficient matrix and B as the right side of AX = B; and to emphasize that X comprises the unknowns, AX = B is referred to as a linear system in X. linear transformation A transformation, say 7\ from a linear space V (of m x /2 matrices) into a linear space W (of p x q matrices) is said to be linear if it satisfies the following two conditions: (1) for all X and Z in V, T(X + Z) = T(K) + HZ); and (2) for every scalar c and for all X in V, T(cX) = cT(X) — in the special case where W = ft, it is customary to refer to a linear transformation from V into W as a linear functional on V. matrix The use of the term matrix is confined (herein) to real matrices, i.e., to rectangular arrarys of real numbers. matrix representation The matrix representation of a linear transformation from
xxiv Some Terminology an /i-dimensional linear space V, with a basis B comprising matrices Vj, V2 V„, into a linear space W, with a basis C comprising matrices Wj, W2 W,„, is the m x n matrix A = [a,j] whose 7th column is (for 7 = 1,2 n) uniquely determined by the equality HVy) = aijY/x + tHjVfz + • • • + amjY/m ; this matrix (which depends on the choice of B and C) is such that if x = [xj } is the n x 1 vector that comprises the coordinates of a matrix V (in V) in terms of the basis B (i.e., V = J^j */Y/).then the m * 1 vector y = {y,} given by the formula y = Ax comprises the coordinates of T(V) in terms of the basis C [i.e., T(\) = £, v,W;]. minimum norm generalized inverse A generalized inverse, say G, of an m x n matrix A is said to be a minimum norm generalized inverse (of A) if (GA)' = GA; or, equivalently, an n x m matrix is a minimum norm generalized inverse of A if it satisfies Moore-Penrose Conditions (1) and (4). Moore-Penrose inverse (and conditions) Corresponding to any m x n matrix A, there is a unique n x m matrix, say G, such that (1) AGA = A (i.e., G is a generalized inverse of A), (2) GAG = G (i.e., A is a generalized inverse of G), (3) (AG)' = AG (i.e., AG is symmetric), and (4) (GA)' = GA (i.e., GA is symmetric). This matrix is called the Moore-Penrose inverse (or pseudoinverse) of A, and the four conditions that (in combination) define this matrix are referred to as Moore-Penrose (or Penrose) Conditions (1)- (4). negative definite An n x n (symmetric or nonsymmetric) matrix A and the quadratic form x'Ax (in an n x 1 vector x) are (by definition) negative definite if —x'Ax is a positive definite quadratic form (or equivalently if —A is a positive definite matrix)—thus. A and x'Ax are negative definite if x'Ax < 0 for every nonnull x in 1Z". negative or positive pair Any pair of elements of an n x n matrix A = {a^} that do not lie either in the same row or the same column, say a,j and a,>y (where /' ^ / and j' t/= j) is (by definition) either a negative pair or a positive pair: it is a negative pair if one of the elements is located above and to the right of the other, or equivalently if either /' > i and j' < j or /' < 1 and j' > j\ otherwise (if one of the elements is located above and to the left of the other, or equivalently if either /' > / and j' > j or /' < / and / < 7), it is a positive pair—note that whether a pair of elements is a negative pair or a positive pair is completely determined by the elements' relative locations and has nothing to do with whether the numerical values of the elements are positive or negative. negative semidefinite An n x n (symmetric or nonsymmetric) matrix A and the quadratic form x'Ax (in an n x 1 vector x) are (by definition) negative semidefinite if —x'Ax is a positive semidefinite quadratic form (or equivalently if —A is a positive semidefinite matrix)—thus, A and x'Ax are negative semidefinite if they are nonposiiive definite but not negative definite, or equivalently if x'Ax < 0 for every x in 1Z" with equality holding for some
Some Terminology xxv nonnull x. neighborhood A neighborhood of an m x n matrix C is a set of the general form {X e 11"' K" : || X — C || < r}, where r is a positive number called the radius of the neighborhood (and where the norm is the usual norm). nonhomogeneous linear system A linear system whose right side (which is a column vector or more generally a matrix) is nonnull. nonnegative definite An n x n (symmetric or nonsymmetric) matrix A and the quadratic form x'Ax (in an n x 1 vector x) are (by definition) nonnegative definite if x'Ax > 0 for every x in 11". nonpositive definite An n x n (symmetric or nonsymmetric) matrix A and the quadratic form x'Ax (in an n x 1 vector x) are (by definition) nonpositive definite if —x'Ax is a nonnegative definite quadratic form (or equivalently if —A is a nonnegative definite matrix)—thus, A and x'Ax are nonpositive definite if x/Ax < 0 for every x in %". nonnull matrix A matrix having 1 or more nonzero elements. nonsingular A matrix is nonsingular if it has both full row rank and full column rank or equivalently if it is square and its rank equals its order. norm The norm of a matrix A in a linear space V is (A • A)1/2—the use of this term is limited herein to norms defined in terms of an inner product; in the case of a quasi-inner product, (A-A)1/2 is referred to as the quasi norm. normal equations A linear system (or the equations comprising the linear system) of the form X'Xb = X'y (in a p x 1 vector b), where X is an n x p matrix and y an n x 1 vector. null matrix A matrix all of whose elements are 0. null space (of a matrix) The null space of an m x n matrix A is the solution space of the homogeneous linear system Ax = 0 (in an /i-dimensional column vector x). or equivalently is the set {x e II"*1 : Ax = 0}. null space (of a transformation) The null space—also known as the kernel— of a linear transformation T from a linear space V into a linear space W is the set {X e V : T(X) = 0}, which is a subspace of V. one to one A transformation T from a set V into a set W is said to be 1-1 (one to one) if each member of the range of T is the image of only one member of V. onto A transformation T from a set V into a set W is said to be onto if T(V) = W (i.e., if the range of T is all of W), in which case T may be referred to as a transformation from V onto W. open set A set Sofmx n matrices is an open set if every matrix in S is an interior point of S. order A (square) matrix of dimensions n x n is said to be of order n. orthogonal complement The orthogonal complement of a subspace U of a linear space V is the set comprising all matrices in V that are orthogonal to U — note that the orthogonal complement of U depends on V as well as U (and
xxvi Some Terminology also on the choice of inner product). orthogonality of a matrix and a subspace A matrix Y in a linear space V is orthogonal to a subspace U (of V) if Y is orthogonal to every matrix in U. orthogonality of two subspaces A subspace U of a linear space V is orthogonal to a subspace W (of V) if every matrix in U is orthogonal to every matrix inW. orthogonality with respect to a matrix For any n x n symmetric nonnegative definite matrix W, two n x 1 vectors, say x and y, are said to be orthogonal with respect to W if x'Wy = 0; an n x 1 vector, say x, and a subspace, say U> of 1Znxl are said to be orthogonal with respect to W if x'Wy = 0 for every y in U\ and two subspaces, say U and V, of 7£"xl are said to be orthogonal with respect to W if x'Wy = 0 for every x in U and every y in V. orthogonal matrix A (square) matrix A is orthogonal if A'A = AA' = I. orthogonal set A finite set of matrices in a linear space V is orthogonal if the inner product of every pair of matrices in the set equals 0. orthonormal set A finite set of matrices in a linear space V is orthonormal if it is orthogonal and if the norm of every matrix in the set equals 1. /An A12 ... Aic partitioned matrix A partitioned matrix, say A21 A22 ... Mc \Ar\ Ar2 ... Arc/ trix that has (for some positive integers r and c) been subdivided into re sub- matrices A,j (i = 1,2,..., r; j = 1,2,..., c), called blocks, by implicitly superimposing on the matrix r — 1 horizontal lines and c—1 vertical lines (so that all of the blocks in the same "row" of blocks have the same number of rows and all of those in the same "column" of blocks have the same number of columns)—in the special case where c = r, the blocks Aj 1, A22 Arr are referred to as the diagonal blocks (and the other blocks are referred to as the off-diagonal blocks). permutation matrix An n x n permutation matrix is a matrix that is obtainable from the n xn identity matrix by permuting its columns; i.e., a matrix of the form (m*, , m*2 M*n), where u\, 1/2 un are respectively the first, second,..., 72 th columns of l„ and where k\, fo,..., k„ is a permutation of the first n positive integers. positive definite An«xn (symmetric or nonsymmetric) matrix A and the quadratic form x'Ax (in an n x 1 vector x) are (by definition) positive definite if x'Ax > 0 for every nonnull x in TV. positive semidefinite Ann xn (symmetric or nonsymmetric) matrix A and the quadratic form x'Ax (in an n x 1 vector x) are (by definition) positive semidefinite if they are nonnegative definite but not positive definite, or equivalently if x'Ax > 0 for every x in 11" with equality holding for some nonnull x.
Some Terminology xxvii principal submatrix A submatrix of a square matrix is a principal submatrix if it can be obtained by striking out the same rows as columns (so that the /th row is struck out whenever the /th column is struck out, and vice versa); the r x r (principal) submatrix of an n x n matrix obtained by striking out its last n — r rows and columns is referred to as a leading principal submatrix (r = l ii). product (of transformations) The product (or composition) of a transformation, say 7\ from a linear space V into a linear space W and a transformation, say S, from a linear space U into V is the transformation from U into W that assigns to each matrix X in U the matrix T[S(X)] (in W)—the definition of the term product (or composition) extends in a straightforward way to three or more transformations. projection (orthogonal) The projection—also known as the orthogonal projection—of a matrix Y in a linear space V on a subspace U (of V) is the unique matrix, say Z, in U such that Y — Z is orthogonal to U\ in the special case where (for some positive integer n and for some symmetric positive definite matrix W) V = TZ"X l and the inner product is the bilinear form x/Wy, the projection of y (an n x 1 vector) on U is referred to as the projection of y on U with respect to W— this terminology can be extended to a symmetric nonnegative definite matrix W by defining a projection of y on U with respect to W to be any vector z in U such that (y — z) J_w U. projection along a subspace For a linear space V of 772 x n matrices and for subspaces li and W such that U © W = V (essentially disjoint subspaces whose sum is V), the projection of a matrix in V, say the matrix Y, on U along W is (by definition) the (unique) matrix Z in U such that Y - Z 6 W. projection matrix (orthogonal) The projection matrix—also known as the orthogonal projection matrix—for a subspace U of 1ZnX l is the unique (« x n) matrix, say A, such that, for every n x 1 vector y, Ay is the projection (with respect to the usual inner product) of y on U — simply saying that a matrix is a projection matrix means that there is some subspace of 7£"xl for which it is the projection matrix. projection matrix (general orthogonal) The (orthogonal) projection matrix for a subspace U of 7£'ixl with respect to an n x n symmetric positive definite matrix W is the unique (n x n) matrix, say A, such that, for every n x 1 vector y. Ay is the projection of y on U with respect to W — simply saying that a matrix is a projection matrix with respect to W means that there is some subspace of %nxl for which it is the projection matrix with respect to W— more generally, a projection matrix for U with respect to an 77. x n symmetric nonnegative definite matrix W is an (71 x n) matrix, say A, such that, for every n x 1 vector y. Ay is a projection of y on U with respect to W. projection matrix for one subspace along another For subspaces U and W (of H" x l) such that U © W = 1Zn x l (essentially disjoint subspaces whose sum is 7£"xl), the projection matrix for U along W is the (unique) n x n matrix,
xxviii Some Terminology say A, such that for every n x 1 vector y, Ay is the projection of y on U along W. QR decomposition The QR decomposition of a matrix of full column rank, say an m x k matrix A of rank k, is the unique decomposition of the form A = QR, where Q is an m x k matrix whose columns are orthonormal (with respect to the usual inner product) and R is a k x k upper triangular matrix with positive diagonal elements. quadratic form A quadratic form in an n x 1 vector x = (jci,...,jc,,)' is a function of x (defined for x € 11") that, for some n x n matrix A = {a,;}, is expressible as x'Ax = £^ j aijxi*j — the matrix A is called the matrix of the quadratic form and, unless n = 1 or the choice for A is restricted (e.g., to symmetric matrices), is nonunique. range The range of a transformation T from a set V into a set W is the set T(V) (i.e., the image of the domain of 7")—in the special case of a linear transformation from a linear space V into a linear space W, the range T(V) of T is a linear space and is referred to as the range space of T. rank (of a linear transformation) The rank of a linear transformation T from a linear space V into a linear space W is (by definition) the dimension dim[nV)] of the range space T(V) of T. rank (of a matrix) The rank of a matrix A is the dimension of C(A) or equiva- lentlyofft(A). rank additivity Two matrices A and B (of the same size) are said to be rank additive if rank(A + B) = rank(A) + rank(B); more generally, k matrices Aj, Ai Ajt (of the same size) are said to be rank additive if rank(5Zf=i ^/) = 5Zf=i rank(A,) (i.e., if the rank of their sum equals the sum of their ranks). reflexive generalized inverse A generalized inverse, say G, of an m x n matrix A is said to be reflexive if GAG = G; or, equivalently, an n x m matrix is a reflexive generalized inverse of A if it satisfies Moore-Penrose Conditions (l)and(2). restriction If 7 is a linear transformation from a linear space V into a linear space W and if li is a subspace of V, then the transformation, say /?, from li into W defined by R(X) = T(X) (which assigns to each matrix in U the same matrix in W assigned by T) is called the restriction of T to li. right inverse A right inverse of an m x n matrix A is an n x m matrix R such that AR = I,„ — a matrix has a right inverse if and only if it has full row rank. row space The row space of an m x n matrix A is the set whose elements consist of all /z-dimensional row vectors that are expressible as linear combinations of the m rows of A. scalar The term scalar is (herein) used interchangeably with real number. scalar multiple (of a transformation) The scalar multiple of a scalar k and a transformation, say 7\ from a linear space V into a linear space W is the
Some Terminology xxix transformation from V into W that assigns to each matrix X in V the matrix kT(K) (in W). Schur complement In connection with a partitioned matrix A of the form A = /T U\ A /W V\ I V w) °r I U T)'the matrix Q = w ~ VT~U is referred to as the Schur complement of T in A relative to T~ or (especially in a case where Q is invariant to the choice of the generalized inverse T~) simply as the Schur complement of T in A or (in the absence of any ambiguity) even more simply as the Schur complement of T. second-degree polynomial A second-degree polynomial in an n x 1 vector x = (.vi .v,,)' is a function, say /(x), of x that is defined for all x in 1Z" and that, for some scalar c\ some n x 1 vector b = {/?/}, and some n x n matrix V = {i»,y}, is expressible as /(x) = c - 2b'x+x'Vx, or in nonmatrix notation as fix) = c-2 £"=1 fc.v,- + £"=i H"=\ vyxiXj — in the special case where c = 0 and V = 0, fix) = -2b'x, which is a linear form (in x), and in the special case where c = 0 and b = 0, fix) = x'Vx, which is a quadratic form (in x). similar An n x n matrix B is said to be similar to an n x n matrix A if there exists an n x n nonsingular matrix C such that B = C-1 AC or, equivalently, such that CB = AC— if B is similar to A, then A is similar to B. singular A square matrix is singular if its rank is less than its order. singular value decomposition An m x n matrix A of rank /• is expressible as A = p(Do J)Q' = PlD|Q' = £ «** = £ -M - where Q = (qif..., q„) is an n x n orthogonal matrix and Dj = diag(^j, ..., sr) an r x r diagonal matrix such that Q'A'AQ = I J ft h where ^1 sr are (strictly) positive, where Q, = (q, qr), Pj = (p, pr) = AQjDj"1, and, for any m x (//? - r) matrix P2 such that P^Po = 0, P = (Pj, P2), where a\ a* are the distinct values represented among si sr, and where (for j = 1 k) Vj = £{/:*,=«,■} P/qJ; any of these four representations may be referred to as the singular value decomposition of A, and s\ sr are referred to as the singular values of A — 5i sr are the positive square roots of the nonzero eigenvalues of A'A (or equivalently AA'), qj q„ are eigenvectors of A'A, and the columns of P are eigenvectors of AA;. skew-symmetric An 11 x n matrix, say A = {0,7}, is (by definition) skew-symmetric if A' = -A; that is, if ay/ = -a/y for all i and j (or equivalently if an = 0 for i = 1 n and ctyt = -<//, for y'^ / = 1 11). solution A matrix, say X*, is said to be a solution to a linear system AX = B (in X)ifAX* = B. solution set or space The collection of all solutions to a linear system AX = B (in X) is called the solution set of the linear system; in the special case of
XXX Some Terminology a homogeneous linear system AX = 0, the solution set may be called the solution space. span The span of a finite set of matrices (having the same dimensions) is defined as follows: the span of a finite nonempty set {Ai,..., A*} is the set consisting of all matrices that are expressible as linear combinations of Ai,..., A*, and the span of the empty set is the set {0}, whose only element is the null matrix. And, a finite set S of matrices in a linear space V is said to span V ifsp(,S) = V. spectral decomposition An n x n symmetric matrix A is expressible as /i k A = QDtf = £ 4q,q; = £ XjEj , /=1 y=l where d\t • • • * dn are the not-necessarily-distinct eigenvalues of A, qlt..., q„ are orthonormal eigenvectors corresponding to d\y • • •. d„, respectively, Q = (q! q„), D = diag(^i d„), [X\ A*} is the spectrum of A, and (for j = 1,..., k) Ej = J^[i:d,=\•) Qi^!' ^y °ftnese ^^ representations may be referred to as the spectral decomposition of A. spectrum The spectrum of an n x n matrix A is the set whose members are the distinct (different) scalars that are eigenvalues of A. subspace A subspace of a linear space V is a subset of V that is itself a linear space. sum (of sets) The sum of 2 nonempty sets, say U and V, of m xn matrices is the set {A + B : A e U, B € V} comprising every (m x n) matrix that is expressible as the sum of a matrix in U and a matrix in V; more generally, the sum of k sets, say U\ Wjt, of m x n matrices is the set (EtiA/ : A,€Wi A*€^}. sum (of transformations) The sum of two transformations, say T and 5, from a linear space V into a linear space W is the transformation from V into W that assigns to each matrix X in V the matrix T(X) + S(X) (in W)—since the addition of transformations is associative, the definition of the term sum extends in an unambiguous way to three or more transformations. symmetric A matrix, say A, is symmetric if A' = A, or equivalently if it is square and (for every / and j) its ijth element equals its jith element. trace The trace of a (square) matrix is the sum of its diagonal elements. transformation A transformation (also known as a function, operator, map, or mapping), say 7\ from a set V, called the domain, into a set W is a correspondence that assigns to each member X of V a unique member of W; the member of W assigned to X is denoted by the symbol T(X) and is referred to as the image of X, and, for any subset U of V, the set of all members of W that are the images of one or more members of U is denoted by the symbol T(U) and is referred to as the image of U — V and W consist of scalars, row or column vectors, matrices, or other "objects". transpose The transpose of an m x n matrix A is the n x m matrix whose //th
Some Terminology element is the ;/th element of A. union The union of 2 sets, say U and V, of m x n matrices is the set comprising all matrices that belong to either or both of U and V; more generally, the union of k sets, say U\,..., W*, of m x n matrices comprises all matrices that belong to at least one oiU\,...Mk- unit (upper or lower) triangular matrix A unit triangular matrix is a triangular matrix all of whose diagonal elements equal one. U'DU decomposition A U'DU decomposition of a symmetric matrix, say A, is a decomposition of the form A = U'DU, where U is a unit upper triangular matrix and D is a diagonal matrix. Vandermonde matrix A Vandermonde matrix is a matrix of the general form /1 jci .v? ... xrl\ X2 -Vo (where *i, *2,..., xn are arbitrary scalars) vi Xn .vj ... *rv vec The vec of an m x n matrix A = (aj, a2,..., a„) is the /n/i-dimensional /ai\ a2 (column) vector I . I obtained by successively stacking the first, second, w ..., nth columns of A one under the other. vech The vech of an n x n matrix A = [ay] is the n(n 4- l)/2-dimensional /a*\ (column) vector where (for / = 1,2 n) a* = («,-,-, aj+\j, VaiV a„iY is the subvector of the ith column of A obtained by striking out its first i — 1 elements. vec-permutation matrix The mn x mn vec-permutation matrix is the unique permutation matrix, denoted by the symbol Kw/lt such that, for every m x n matrix A, vec(A') = Km/Jvec(A) — the vec-permutation matrix is also known as the commutation matrix. zero transformation The linear transformation from a linear space V into a linear space W that assigns to every matrix in V the null matrix (in W) is called the zero transformation.
1 Matrices EXERCISE 1. Show that, for any matrices A, B, and C (of the same dimensions), (A + B) + C = (C + A)+B. Solution. Since matrix addition is commutative and associative, (A + B) + C = C + (A + B) = (C + A)+B. EXERCISE 2. For any scalars c and k and any matrix A, c{kk) = (ck)\ = (kc)A = A-(cA), (*) and, for any scalar c, m x n matrix A, and n x p matrix B, cAB = (cA)B = A(cB). (**) Using results (*) and (**) (or other means), show that, for any m x n matrix A and n x p matrix B and for arbitrary scalars c and A\ (cA)(kB) = (cA-)AB. Solution. Making use of results (**) and (*), we find that (cA)(*B) = *(cA)B = k(cAB) = (ck)AB. EXERCISE 3. (a) Verify the associativeness of matrix multiplication; that is, show that, for any m x n matrix A = {fl/;}, n x q matrix B = {&/*}. and q x r matrix C = [cksl A(BC) = (AB)C
2 1. Matrices (b) Verify the distributiveness with respect to addition of matrix multiplication; that is, show that, for any m x n matrix A = {ay} and n xq matrices B = [bjk] and C = [cjkl A(B + C) = AB + AC. Solution, (a) The jsth element of BC equals J^k bjkCks, and similarly the ikth element of AB equals J\- aybjt. Thus, the /5th element of A(BC) equals ^aU\Y,bS*cks) = Jl\J2aUbJkCksJ = J2 \J2aUbJkck') = Y, [12aiJbJk)Cks' and 52*(52# a'ijbjk)cks equals the /5th element of (AB)C. Since each element of A(BC) equals the corresponding element of (AB)C, we conclude that A(BC) = (AB)C. (b) Observing that the jkth element of B + C equals bjk + cj*. we find that the /A-th element of A(B + C) equals £aij (bjk + Cjk) = £ {oijbjk + aijcjt) = £aybjk + £aucJk. j J J J Further, observing that J2 • ay bjk is the ikth element of AB and that J^j aycjk is the ikth element of AC, we find that J2/ aubjk + 52y aUcjk equals the ikth element of AB + AC. Since each element of A(B + C) equals the corresponding element of AB + BC, we conclude that A(B + C) = AB + BC. EXERCISE 4. Let A = {ay} represent an m x n matrix and B = {by} apxm matrix. (a) Let x = {.\/} represent an n-dimensional column vector. Show that the /th element of the p-dimensional column vector BAx is in n Y,bijHaJkXk- (E1) y=I *-l (b) Let X = {xy} represent an n x q matrix. Generalize formula (E.1) by expressing the /Vth element of the p x q matrix BAX in terms of the elements of A, B, and X. (c) Let x = {.y,} represent an ?i-dimensional column vector and C = {cy} a q x /; matrix. Generalize formula (E.1) by expressing the /th element of the <y-dimensional column vector CBAx in terms of the elements of A, B, C, and x. (d) Let y = {v/| represent a p-dimcnsional column vector. Express the /th element of the //-dimensional row vector y'BA in terms of the elements of A, B, andy.
1. Matrices 3 Solution, (a) The jth element of the vector Ax is £J!=1 ajkxk. Thus, upon regarding BAx as the product of B and Ax, we find that the /th element of BAx is (b) The /rth element of BAX is m n J^bijJ^ajkxkr, y=I Jt=I as is evident from Part (a) upon regarding the /rth element of BAX as the /th element of the product of BA and the rth column of X. (c) According to Part (a), the 5th element of the vector BAx is in 11 j=i k=i Thus, upon regarding CBAx as the product of C and BAx, we find that the /th element of CBAx is P in n Y,CisJ^b'jJlaJkXk- s=\ j=\ k=\ (d) The /th element of the row vector y'BA is the same as the /th element of the column vector (y'BA)' = A'B'y. Thus, according to Part (a), the /th element of y'BA is //; p Y,aJiJlbkjyk- j=i k=i EXERCISE 5. Let A and B represent n x n matrices. Show that (A +B)(A -B) = A2-B2 if and only if A and B commute. Solution. Clearly, (A + B)(A-B)=A(A-B) + B(A-B) = A2-AB + BA-B2. Thus, (A + B)(A-B) = A2-B2 if and only if -AB + BA = 0 or equivalently if and only if AB = BA (i.e., if and only if A and B commute). EXERCISE 6. (a) Show that the product AB of two n x n symmetric matrices A and B is itself symmetric if and only if A and B commute. (b) Give an example of two symmetric matrices (of the same order) whose product is not symmetric.
4 1. Matrices Solution, (a) Since A and B are symmetric, (AB)' = B'A' = BA. Thus, if AB is symmetric, that is, if AB = (AB)', then AB = BA, that is, A and B commute. Conversely, if AB = BA, then AB = (AB)'. (b)TakeA=(* MandB=(° M.Then, —(SJ)'(IO-"- EXERCISE 7. Verify (a) that the transpose of an upper triangular matrix is lower triangular and (b) that the sum of two upper triangular matrices (of the same order) is upper triangular. Solution. Let A = [ay] represent an upper triangular matrix of order n. Then, by definition, the ijth element of A' is ay,-. Since A is upper triangular, ajt = 0 for / < j = 1 n or equivalently for j > i = 1 «. Thus, A' is lower triangular, which verifies Part (a). Let B = {by} represent another upper triangular matrix of order n. Then, by definition, the ijth element of A + B is ay + by. Since both A and B are upper triangular, ay = 0 and by = 0 for j < i = 1 /z, and hence ay + by = 0 for j < i = 1 n. Thus, A + B is upper triangular, which verifies Part (b). EXERCISE 8. Let A = [ay] represent an n x n upper triangular matrix, and suppose that the diagonal elements of A equal zero (i.e., that a\ \ = #22 = • • • = a„„ = 0). Further, let p represent an arbitrary positive integer. (a) Show that, for / = 1 n and j = 1,..., min(/*, i + p — 1), the ijth element of \p equals zero. (b) Show that, for i > n — p + 1, the ith row of \p is null. (c) Show that, for p > iu Ap = 0. Solution. For /, k = 1 /1, let bit represent that /fcth element of \p. (a) The proof is by mathematical induction. Clearly, for / = 1 n and j — 1,..., min(/2, / +1 — 1), the ijth element of A1 equals zero. Now, suppose that, for / = 1 n and j = 1 min(/7, i+p — 1), the ijth element of A*7 equals zero. Then, to complete the induction argument, it suffices to show that, for / = 1 n and j = 1 min(/2, i + p), the ijth element of Ap+1 equals zero. Observing that \p+l = A^A, we find that, for / = 1 n and j = 1 min(«, i + p), the ijth element of Ap+1 equals 11 min(i/,/+p-l) ,1 ^bikakj = ]T 0atj+ ^2 bikakJ k=\ k=l k=i+p (where, if/ > n — 77, the sum J2'k=i+P ^ikOkj is degenerate and is to be interpreted as 0)
1. Matrices 5 = 0 (since, for k > j, cikj = 0). (b)For/ > rt-p+l,min(/2, i+p— 1) = n (since/ > n—p+1 <£► i+p— 1 > n). Thus, for i >n — p + 1, it follows from Part (a) that all n elements of the /th row of Ap equal zero and hence that the /th row of \p is null. (c) Clearly, forp>nt/i — p+l<l. Thus, for p > «, it follows from Part (b) that all // rows of \p are null and hence that \p = 0.
Submatrices and Partitioned Matrices EXERCISE 1. Let A* represent anrxj submatrix of an m x n matrix A obtained by striking out the i\ /,„-rth rows and j\ ;„-jth columns (of A), and let B* represent the s x r submatrix of A' obtained by striking out the j\ y,/-5th rows and u,..., /„;-,th columns (of A'). Verify that b* = a;. Solution. Let i* i* (/*<••• < i*) represent those r of the first m positive integers that are not represented in the sequence /'i,..., /,„_r. Similarly, let j*,..., j* (j* < ••• < j*) represent those s of the first n positive integers that are not represented in the sequence j\ jn-s. Denote by aij and fc/y the iyth elements of A and A;, respectively. Then, a: =: wr ••• fl|;/r flW bnn = B* EXERCISE 2. Verify (a) that a principal submatrix of a symmetric matrix is symmetric, (b) that a principal submatrix of a diagonal matrix is diagonal, and (c) that a principal submatrix of an upper triangular matrix is upper triangular.
8 2. Submatrices and Partitioned Matrices Solution. Let B = {bij} represent the rxr principal submatrix of an n x n matrix A = {aij} obtained by striking out all of the rows and columns except the fci, fo. ..., krth rows and columns (where k\ < k% < • • • < kr). Then, by = aktkj (i. j = 1 r). (a) Suppose that A is symmetric. Then, for /, j = 1 r, by = a^kj = Gkjk, = bji. (b) Suppose that A is diagonal. Then, for j ^ / = 1 r, fc/y = a*,*, = 0. (c) Suppose that A is upper triangular. Then, for j < i = 1 r, by = ak,k} = 0. EXERCISE 3. Let /An A12 0 A22 Vo 0 Alr\ A2r Arr>/ represent an«x« upper block-triangular matrix whose ij\h block A,-y is of dimensions n/ xrij (j >i = 1 r). Show that A is upper triangular if and only if each of its diagonal blocks An, A22 Arr is upper triangular. Solution. Let ats represent the tsth element of A (/, s = 1,..., n). Then, /«/il+-+ii;_i+l,ii|+-+iiy_l + l ••• tfii|+-+«,-i + l.fl|+-+/i;-l+rt/ \ ^111+-+11/-1+/1,-./11+-+/1^-1+1 • • • tfiii+-+M/_i+/j,-,ij|+—+ii,_i+rt,/ 0*>' = 1 r). Suppose that A is upper triangular. Then, by definition, ats = 0 for s < t = 1 n. Thus, anx+...+ni_x+k,nx+...+n,-\+i (which is the kith element of the /th diagonal block A,-,-) equals zero for I < k = 1 h,\ implying that A,-,- is upper block-triangular (/ = 1 r). Conversely, suppose that Ai 1, A 22» • • •» A/-/- are upper triangular. Let t and s represent any integers (between 1 and /z, inclusive) such that ats ^ 0. Then, clearly, for some integers i and j > /, ats is an element of the submatrix A/y\ say the A7th element, in which case t = n \ -\ h/i/_ 1 + k and s = n \ -\ \-»j-1 +/. If j > /, then (since k < /z,-) t < s. Moreover, if j = /, then (since A,-,- is upper triangular) k < /, implying that t < s. Thus, in either case, t < s. We conclude that A is upper triangular. EXERCISER Let A = /A„ A21 A12 A22 .. A,c\ •• A2, VAri Ar2
2. Submatrices and Partitioned Matrices 9 represent a partitioned m x n matrix whose //th block A/y is of dimensions /w/ x /zy. Verify that /A'n Ai, ... a;a A' = A' A' 12 72 K rl W A'2r ... a; 're/ in other words, verify that A' can be expressed as a partitioned matrix, comprising c rows and r columns of blocks, the //th of which is the transpose Ay, of the y/th block Ay/ of A. And, letting B = /Bn B,2 B21 B22 \BMi B„2 B!v\ B2v B, >uv/ represent a partitioned p x q matrix whose /yth block B,y is of dimensions p/ x #y, verify also that if c = u and w* = pk (k = 1,..., c) [in which case all of the products A/frBjt/ (/ = 1 r;j = l v;k = \ c), as well as the product AB, exist], then /Fn F12 ... Fi„\ I F21 F22 ... $2v AB = . \Fri Fr2 ... FrvJ where F/y = ££=1 A/*B*y = A/iBiy + A/2B2y + • • ■ + A/cBcy. Solution. Let «/y, fc/y, /z/y, and sy represent the //th elements of A, B, A', and AB, respectively. Define H/y to be a matrix of dimensions /i/ x mj (i = /,..., c; j = 1,.... r) such that A' = /Hn H12 H21 H22 VHcl Hc2 Clearly, H/y is the submatrix of A' obtained by striking out the first nH h«/-i and last w+i + • • • + nc rows of A' and the first /711 + ---+ wy-i and last ntj+i + • • • + mr columns of A'; and Ay/ is the submatrix of A obtained by striking out the first mi H h my_ 1 and last /ny+H h mr rows of A and the first n 1 H h «/-1 and last «/+1H Vnc columns of A. Thus, it follows from result (1.1) that H/y=A}/.
10 2. Submatrices and Partitioned Matrices Further, define S/y to be a matrix of dimensions m{ x qj (i = 1,..., n j = 1 v) such that /Su S12 ... SiA I S21 S22 • • • Slv AB = . \Srl Sr2 ... Srv/ Then, for w = 1 /«,• and c = 1 #y, the u>zth element of S/y is Smi+—+in,-i+w,qi+'"+<Jj-i+Z 111+—+»<• = ^ Am,+...+01,--1 +u>,£ ^£,91+».+^_j+z f=I C fl I+•••+!!*-1+*£ = 2_^ ^ «iilj+.+l«,-|+lu.£ h.qi+-+qj-i+Z *=1 f=il|+-+«A_l + I C ilk = / ,/ ,fli*i+-+iWf-|+w.W|+-+fljt-|+> ^il|+- +HA-I+/,91+"H-^-i+C- *=I /=1 And, upon observing that a,,,,+...+,„, ..,+^.,,,+...+,,^,+, is the utfth element of A,* and that £,,,+...+,,^,+,^,+...+^,+- is the tzth element of B*y, it is clear that / ,am1 +-+111,-1 +u>.ni+-+Wfc-i +t bni+...+,tk_l+t.ql+...+qj_i+z is the u>zth element of A^B*/ and hence that s,„,+...+„,._,+,y.9,+...+^_,+- is the wzth element of F,y. Thus,
3 Linear Dependence and Independence EXERCISE 1. For what values of the scalar k are the three row vectors (k, 1,0), (1, £, 1), and (0, 1, k) linearly dependent, and for what values are they linearly independent? Describe your reasoning. Solution. Let x\* A2, and .*3 represent any scalars such that xi(k, 1,0)+a*2(1,A', 1) + a-3(0, 1,*) = 0, or equivalently such that x\k + .\2 = 0, x\ + xik + A3 = 0, Xo + A3A = 0, or also equivalently such that X2 = -kx3 = -kxi. (S.l) kx2 = -xi-x3. (S.2) Suppose that k = 0. Then, conditions (S.l) and (S.2) are equivalent to the conditions A2 = 0 and X3 = — x\. Alternatively, suppose that k ^ 0. Then, conditions (S.l) and (S.2) imply that A3 = a*i and — k2x\ = kxo = —2x\ and hence that k2 = 2 or A3 = a*2 = ai = 0. Moreover, if k2 = 2, then either k = >/2, in which case conditions (S.l) and (S.2) are equivalent to the conditions A3 = aj and .V2 = — >/2.vj, or k = —y/l% in which case conditions (S. 1) and (S.2) are equivalent to the conditions A3 = aj and .v2 = v^2.vi.
12 3. Linear Dependence and Independence Thus, there exist values of *i, *2, and *3 other than x\ = X2 = x$ = 0 if and only if k = 0 or k = ±y/2. And, we conclude that the three vectors (k, 1,0), (1, k, 1), and (0,1, k) are linearly dependent if/: = 0 or k = ±V2, and linearly independent, otherwise. EXERCISE 2. Let A, B, and C represent three linearly independent m x n matrices. Determine whether or not the three pairwise sums A + B, A + C, and B + C are linearly independent. [Hint. Take advantage of the following general result on the linear dependence or independence of linear combinations: Letting Ai, A2,..., A* represent m x n matrices and for j = 1,..., r, taking C/ = *iyAi + X2jA2 -\ 1- Xkj\k (where x\j, xoj,..., xy are scalars) and letting xj = (x\j,X2j **/)'» the linear combinations Ci, C2,..., Cr are linearly independent if Ai, A2 A* are linearly independent and xi,X2 xr are linearly independent, and they are linearly dependent if xi, X2 xr are linearly dependent.] Solution. It follows from the result cited in the hint that A + B, A + C, and B + C are linearly independent if (and only if) the three vectors (1,1,0)\ (1,0,1/, and (0,1,1,)' are linearly independent. Moreover, for any scalars jq, .yi, and *3 such that x,(l, l,0)' + .v2(l,0, l)'+*3(0,1, 1)' = 0, we have that *i + .Y2 = *i + .Y3 = X2 4- .V3 = 0, implying that 2v3 = 0 + 2.Y3 = (JCi +X2) +2*3 = (XI +*3) + fo + JT3) = 0 +0 = 0 and x\ = X2 = — V3 and hence that jq = 0 and jci = .Y2 = 0. Thus, (1,1,0)\ (1,0,1)', and (0,1,1)' are linearly independent. And, we conclude that A+B, A+ C, and B + C are linearly independent.
4 Linear Spaces: Row and Column Spaces EXERCISE 1. Which of the following two sets are linear spaces: (a) the set of all n x n upper triangular matrices; (b) the set of all n x n nonsymmetric matrices? Solution. Clearly, the sum of two nxn upper triangular matrices is upper triangular. And, the matrix obtained by multiplying any n x n upper triangular matrix by any scalar is upper triangular. However, the sum of two n x n nonsymmetric matrices is not necessarily nonsymmetric. For example, if A is an n x n nonsymmetric matrix, then —A and A' are nonsymmetric, yet the sums A + (—A) = 0 and A + A' are symmetric. Also, the product of the scalar 0 and any nxn matrix is the null matrix, which is symmetric. Thus, the set of all n x n upper triangular matrices is a linear space, but the set of all n x n nonsymmetric matrices is not. EXERCISE 2. Letting A represent anmxn matrix and B an m x p matrix, verify that (1) C(A) C C(B) if and only if ft(A') c ft(B'), and (2) C(A) = C(B) if and only if ft(A') = ft(B'). Solution. (1) Suppose that ft(A') C 1KB'). Then, for any vector x in C(A), we have (in light of Lemma 4.1.1) that x7 € ft(A'), implying that x' € TZ(B') and hence (in light of Lemma 4.1.1) that x € C(B). Thus, C(A) c C(B). Conversely, suppose that C(A) c C(B). Then, for any m-dimensional column vector x such that x; € 7£(A;), we have that x € C(A), implying that x € C(B) and hence thatx/ € ft(B'). Thus, ft(A') c ft(B'). We conclude that C(A) C C(B) if and only if ft(A') C ft(B'). An alternative verification of Part (1) is obtained by taking advantage of Lemma
14 4. Linear Spaces: Row and Column Spaces 4.2.2. We have that C(A)cC(B) & A = BK for some matrix K <e> A' = K'B' for some matrix K <& ft(A') C ft(B'). (2) If ft(A') = ft(B'), then ft(A') C ft(B') and ft(B') C ft(A'), implying [in light of Part (1)] that C(A) C C(B) and C(B) C C(A) and hence that C(A) = C(B). Similarly, if C(A) = C(B), then C(A) C C(B) and C(B) C C(A), implying that ft(A') C ft(B') and ft(B') C ft(A') and hence that ft(A') = ft(B'). Thus, C(A) = C(B) if and only if ft (A') = ft(B'). EXERCISE 3. Let U and W represent subspaces of a linear space V. Show that if every matrix in V belongs to U or W, then U = V or W = V. Solution. Suppose that every matrix in V belongs to U or W. And, assume (for purposes of establishing a contradiction) that neither U = V nor W = V. Then, there exist matrices A and B in V such that A ¢. U and B ¢. W. And, since A and B each belong to U or W, A 6 W and B e U. Clearly,A = B-(B-A)andB = A+(B-A),andB-AeWorB-AeW. If B - A e W, then B - (B - A) e U and hence A e U. If B - A e W, then A + (B — A) e W and hence B e W. In either case, we arrive at a contradiction. We conclude that U = V or W = V. EXERCISE 4. Let A, B, and C represent three matrices (having the same dimensions) such that A + B + C = 0. Show that sp(A, B) = sp(A, C). Solution. Let E represent an arbitrary matrix in sp(A, B). Then, E = d\ + AB for some scalars d and fc, implying (since B = —A — C) that E = d\ + k(-k -C) = (d- k)\ + (-k)C e sp(A, C). Thus, sp(A, B) C sp(A, C). And, it follows from an analogous argument that sp(A, C) C sp(A, B). We conclude that sp(A, B) = sp(A, C). EXERCISE 5. Let Aj A* represent any matrices in a linear space V. Show that sp(Ai Ajt) is a subspace of V and that, among all subspaces of V that contain A| A*, it is the smallest [in the sense that, for any subspace U (of V) that contains A|,..., A*, sp(Ai AjlcW]. Solution. Let U represent any subspace of V that contains Ai A*. It suffices (since V itself is a subspace of V) to show that sp(A| A*) is a subspace of U. Let A represent an arbitrary matrix in sp(A| A*). Then, A = .vi A| -\ h .vjtAjt for some scalars .vi a*, implying that A e U. Thus, sp(A| A*) is a subset of U, and, since sp(A|,..., A*) is a linear space, it follows that sp(A| Ajt) is a subspace of U.
4. Linear Spaces: Row and Column Spaces 15 EXERCISE 6. Let Aj,..., Ap and Bj Bg represent matrices in a linear space V. Show that if the set {Ai A,,} spans V, then so does the set {Aj Ap, Bi B,,}. Show also that if the set {Aj Ap, Bi Bg} spans V and if Bj B^ are expressible as linear combinations of A| A7„ then the set {Ai Ap} spans V. Solution. It suffices (as observed in Section 4.3c) to show that if Bj Bg are expressible as linear combinations of Ai Ap, then any linear combination of the matrices Aj Ar, Bj Bg is expressible as a linear combination of Ai,..., Ap and vice versa. Suppose then that there exist scalars k\j kpj such that By- = 52f kij^i U = 1 q)- Then, for any scalars x\ xp,y\ yq, 5>A, + £ yjBj = ]>>, + £ yjku)Ah i J ' J which verifies that any linear combination of Aj,..., Ap, Bj Bg is expressible as a linear combination of Aj Ap. That any linear combination of Ai Ap is expressible as a linear combination of Ai Ap, Bj B^ is obvious. EXERCISE 7. Suppose that {Aj A*} is a set of matrices that spans a linear space V but is not a basis for V. Show that, for any matrix A in V, the representation of A in terms of Aj Ajt is nonunique. Solution. Let x\ .y* represent any scalars such that A = 52/=i .v,A/. [Since sp(Ai Ajt) = V, such scalars necessarily exist.] Since the set {Aj A*} spans V but is not a basis for V, it is linearly dependent and hence there exist scalars z\ Zk* not all zero, such that £/=1 7,A,- = 0. Letting y,- = jc,- + Zi (i = 1 A), we obtain a representation A = J^=\ yi^i different from the representation A = 52/-1 -v/A/. EXERCISER Let A = (a) Show that each of the two column vectors (2, -1, 3, -4)' and (0,9, -3, 12/ is expressible as a linear combination of the columns of A [and hence is in C(A)]. (b) A basis, say 5*, for a linear space V can be obtained from any finite set S that spans V by successively applying to each of the matrices in S the following algorithm: include the matrix in S* if it is nonnull and if it is not expressible as a linear combination of the matrices already included in S*. Use this algorithm to find a basis for C( A). (In applying the algorithm, take the spanning set S to be the set consisting of the columns of A.) (c) What is the value of rank(A)? Explain your reasoning. /0 0 0 \o 1 -2 2 -4 0 0 2 -2 -3 6 5 1 2" 2 2 o
16 4. Linear Spaces: Row and Column Spaces (d) A basis for a linear space V that includes a specified set, say 7\ of r linearly independent matrices in V can be obtained by applying the algorithm described in Part (b) to the set S whose first r elements are the elements of T and whose remaining elements are the elements of any finite set U that spans V. Use this generalization of the procedure from Part (b) to find a basis for C(A) that includes the two column vectors from Part (a). (In applying the generalized procedure, take the spanning set U to be the set consisting of the columns of A.) Solution, (a) Clearly, and / \ 0\ 9 -3 12/ 1 -2 2 -4 + (1/2) /2\ 2 2 = (-3) + (3/2) (2 (b) The basis obtained by applying the algorithm comprises the following 3 vectors: 0\ 0 2 "2/ /2\ 2 2 (c) Rank A = 3. The number of vectors in a basis for C(A) equals 3 [as is evident from Part (b)], implying that the column rank of A equals 3. (d) The basis obtained by applying the generalized procedure comprises the following 3 vectors: / 0\ 0 2 -2/ EXERCISE 9. Let A represent a <y x p matrix, B a /? x /? matrix, and C an //i x # matrix. Show that (a) if rank(CAB) = rank(C), then rank(CA) = rank(C) and (b) if rank(CAB) = rank(B), then rank(AB) = rank(B). Solution, (a) Suppose that rank(CAB) = rank(C). Then, it follows from Corollary 4.4.5 that rank(C) > rank(CA) > rank(CAB) = rank(C) and hence that rank(CA) = rank(C).
4. Linear Spaces: Row and Column Spaces 17 (b) Similarly, suppose that rank(CAB) = rank(B). Then, it follows from Corollary 4.4.5 that rank(B) > rank(AB) > rank(CAB) = rank(B) and hence that rank(AB) = rank(B). EXERCISE 10. Let A represent anmx/i matrix of rank r. Show that A can be expressed as the sum of/* matrices of rank 1. Solution. According to Theorem 4.4.8, there exist an m x r matrix B and an r x n matrix T such that A = BT. Let bj br represent the first rth columns of B and t'j t'r the first rth rows of T. Then, applying formula (2.2.9), we find that A = £Ay, where (for j = 1 r) A; = byt'.. Moreover, according to Theorem 4.4.8, rank(B) = rank(T) = r, and it follows that bj br and t'j tj. are nonnull and hence that Aj Ar are nonnull. And, upon observing (in light of Corollary 4.4.5 and Lemma 4.4.3) that rank(Ay) < rank(by) < 1, it is clear that rank(Ay) = 10 = 1 r). EXERCISE 11. Let A represent an m x n matrix and C a q xn matrix. (a) Confirm that U(C) = ll(^) & 11(A) C 11(C). (b) Confirm that rank(C) < rank(r L with equality holding if and only if 11(A) C 11(C). Solution, (a) Suppose that 11(A) C 11(C). Then, according to Lemma 4.2.2, there exists an m x q matrix L such that A = LC and hence such that I c J = I - JC. Thus, 1l(^\ C 11(C), implying [since 11(C) C K[q\] ™at W) = ft(£Y Conversely, suppose that 11(C) = ^(c)- Then, since 11(A) C n[c) • 11(A) C 11(C). Thus, we have established that 11(C) = TilQ J & 11(A) C 11(C). (b) Since (according to Lemma 4.5.1) 11(C) C 7e(cj ,itfollows from Theorem 4.4.4 that rank(C) < rank(cj. Moreover, if 11(A) C ft(C), then [according
18 4. Linear Spaces: Row and Column Spaces to Part (a) or to Lemma 4.5.1] 11(C) = ft(£J and consequently rank(C) = rankf c J. And, conversely, if rank(C) = rankf c V then since 11(C) C u[ c ) , it follows from Theorem 4.4.6 that 11(C) = 111 c ) and hence [in light of Part (a) or of Lemma 4.5.1] that ft (A) c 11(C). Thus, rank(C) < rankf r) , with equality holding if and only if ft(A) C 11(C).
5 Trace of a (Square) Matrix EXERCISE 1. Show that for any m x n matrix A,nxp matrix B, and p x q matrix C, tr(ABC) = tr(B'A'C) = tr(A'C'B'). Solution. Making use of results (2.9) and (1.5), we find that tr(ABC) = tr(CAB) = tr[(CAB)'] = tr(B'A'C) = tr(A'C'B'). EXERCISE 2. Let A, B, and C represent n x n matrices. (a) Using the result of Exercise 1 (or otherwise), show that if A, B, and C are symmetric, then tr(ABC) = tr(BAC). (b) Show that [aside from special cases like that considered in Part (a)] tr(BAC) is not necessarily equal to tr(ABC). Solution, (a) If A, B, and C are symmetric, then B'A'C = BAC and it follows from the result of Exercise 1 that tr(ABC) = tr(BAC). (b) Let A = diag(A*, 0), B = diag(B+, 0), and C = diag(C*, 0), where A* = (o o)' B* = (-i o} c* = (o -l} Then, A*B*C* = 0 and B*A+C* = (_, _- ), and, observing that ABC = diag(A*B*C*, 0) and BAC = diag(B+A*C+, 0) and making use of result (1.7),
20 5. Trace of a (Square) Matrix we find that tr(BAC) = tr(B*A+C*) = 2 # 0 = tr(A*B*C*) = tr(ABC). EXERCISE 3. Let A represent an /2 x n matrix such that A'A = A2. (a) Show that tr[(A - A')'(A - A')] = 0. (b) Show that A is symmetric. Solution, (a) Making use of results (2.3) and (1.5), we find that tr[(A - A')' (A - A')] = tr[A'A - A'A' - AA + AA'] = tr(A'A) - tr[(AA)'] - tr(A2) + tr(AA') = tr(AA') - tr[(AA)'] = tr(A'A) - tr[(AA)'] = tr(A'A) - tr(AA) = 0. (b) In light of Lemma 5.3.1, it follows from Part (a) that A - A' = 0 or equivalent^ that A' = A.
6 Geometrical Considerations EXERCISE 1. Use the Schwarz inequality to show that, for any two matrices A and B in a linear space V, ||A + B||<||A|| + ||B||, with equality holding if and only if B = 0 or A = A'B for some nonnegative scalar k. (This inequality is known as the triangle inequality.) Solution. We have that ||A + B||2 = (A + B)-(A + B) = ||A||2+2(A-B) + ||B||2 <||A||2+2|A-B|+ ||B||2 <||A||2+2||A||||B||+ ||B||2 (using the Schwarz inequality) = (HA|| + ||B||)2 or equivalent^ that (S.l) (S.2) |A + B|<||A| + ||B||. For this inequality to hold as an equality, it is necessary and sufficient that both of inequalities (S.l) and (S.2) hold as equalities. Recalling (from, for instance, Theorem 6.3.1) the conditions under which the Schwarz inequality holds as an equality, we find that inequalities (S.l) and (S.2) both hold as equalities if and only if B = 0 or A = kB with k > 0.
22 6. Geometrical Considerations EXERCISE 2. Letting A, B, and C represent arbitrary matrices in a linear space V, show that (a) 8(B, A) = 8(A, B), that is the distance between B and A is the same as that between A and B; (b) a(A,B)>0, ifA#B, = 0, ifA = B, that is, the distance between any two matrices is greater than zero, unless the two matrices are identical, in which case the distance between them is zero; (c) 3(A, B) < 8(A, C)+ 8(C, B), that is, the distance between A and B is less than or equal to the sum of the distances between A and C and between C and B; (d) 8(\% B) = 8(A + C, B + C), that is, distance is unaffected by a translation of "axes." [For Part (c), use the result of Exercise 1, i.e., the triangle inequality.] Solution, (a) «(B.A) = ||B-A|| = ||(-1)(A-B)|| = |-1|||A-B|| = ||A-B||=3(A,B). (b) 3(A, B) = || A - B || > 0, if A - B ^ 0 or equivalently if A # B, = 0, if A - B = 0 or equivalently if A = B. (c) S(A,B) = IIA-BH = ||(A-C) + (C-B)|| < ||A-C|| + ||C-B|| = a(A,C)+S(C,B). (d) a<A + GB + C) = ||(A + C)-(B + C)| = |A-B|=a<A,B). EXERCISE 3. Let wjy,, and W3 represent the three linearly independent 4- dimensional row vectors (6,0, -2,3). (-2.4,4.2), and (0,5. -1,2), respectively, in the linear space ft4, and adopt the usual definition of inner product. (a) Use Gram-Schmidt orthogonalization to find an orthonormal basis for the linear space spfw^, w',, w',).
6. Geometrical Considerations 23 (b) Find an orthonormal basis for ft4 that includes the three orthonormal vectors from Part (a). Do so by extending the results of the Gram-Schmidt orthogonaliza- tion [from Part (a)] to a fourth linearly independent row vector such as (0, 1, 0, 0). Solution, (a) The 3 orthogonal vectors obtained by applying the formulas (for Gram-Schmidt orthogonalization) of Theorem 6.4.1 are: /, =w', =(6,0,-2,3), /2 = w; - (-2/7)/, = (1/7)(-2,28,24,20), yi = w3 -(13/21)/, - (8/49)/, = (1/147)(-118,371, -411, -38). By normalizing y',, y2, and y3, we obtain a basis for sp(w',, w2, w3) consisting of the following 3 vectors: z', =(1/7)/, =(1/7)(6,0,-2,3), z2 = (1/%; = (1/42)(-2,28,24,20), z3 = (21609/321930)1/2y3 = (321930)_1/2(—118,371, -411, -38). (b) An orthonormal basis for ft4 can be obtained by extending the results of the Gram-Schmidt orthogonalization to a fourth linearly independent vector W4. Taking w^ = (0,1,0,0) and applying the formulas of Theorem 6.4.1, we obtain the following vector, which is orthogonal to y',, y2, and y3: /4 = K ~ (371/2190)/3 - (1/9)/, - (0)/, = (1/321930)(53998,41209,29841, -88102). The set consisting of the normalized vector z'4 = [321930/(13266413370)l/2]yi = (13266413370r1/2(53998,41209,29841, -88102), together with z\, z'ly and z3, is an orthonormal basis for ft4. EXERCISE 4. Let {Aj,..., A*} represent a nonempty (possibly linearly dependent) set of matrices in a linear space V. (a) Generalize the results underlying Gram-Schmidt orthogonalization (which are for the special case where the set {Aj A*} is linearly independent) by showing (1) that there exist scalars Xjj (i < j = 1 k) such that the set comprising the k matrices Bi=Aj, B2 = A2-A12B1, B7- = Ay - Xj-ijhj-\ A-jjBj, B^ = A* - **_ijtBjt-i a-u-Bi
24 6. Geometrical Considerations is orthogonal; (2) that, for j = 1 k and for those i < ; such that B, is nonnull, xij is given uniquely by A;-B, *;=B^r and (3) that the number of nonnull matrices among Bj B& equals dim[sp(Aj, ...,A*)]. (b) Describe a procedure for constructing an orthonormal basis for sp(Ai,..., Ait). Solution, (a) The proof of (1) and (2) is by mathematical induction. Asseruons (1) and (2) are clearly true for k = 1. Suppose now that they are true for a set of k — 1 matrices. Then, there exist scalars x,j (/ < j = 1 k — \) such that the set comprising the k - 1 matrices Bj Bjt_i is orthogonal, and, for 7 = 1 k-\ and for those i < j such that B, is nonnull, xjj is given uniquely by _A,-B, ^"b^bT Moreover, for / = 1,..., k — 1, we find (as in the proof of the results underlying Gram-Schmidt orthogonalization in Theorem 6.4.1) that B* -B, = 0 if and only if A* -8,-^(8,-8,)=0. For those / (between 1 and k — 1) such that 8,- = 0, this equation is satisfied by any *,*, and, for those i such that B,- £ 0, it has the unique solution _A*-BI This completes the induction argument, thereby establishing (1) and (2). Consider now Assertion (3). Each of the matrices Bj,..., B* can (by repeated substitution) be expressed as a linear combination of Aj A*. Conversely, each of the matrices A j A* can be expressed as a linear combination of Bj B*. Thus, sp(Bj B*) = sp(Aj A*). Since the set {Bj B#} is orthogonal, we conclude — in light of Lemma 6.2.1 and Theorem 4.3.2 — that the nonnull matrices among B| B* form a basis for sp(Aj A*) and hence that the number of such matrices equals dim[sp(Aj A*)]. (b) An orthonormal basis for sp(Aj A*) can be constructed by making use of the formulas for Bj B* from Part (a). The basis consists of those matrices obtained by normalizing the nonnull matrices among Bj B*. EXERCISE 5. Let A represent an m x k matrix of rank r (where r is possibly less than k). Generalize the so-called QR decomposition of A, which is for the special case where r = k and is obtainable through the application of Gram-Schmidt orthogonalization to the columns of A. Do so by using the results of Exercise 4 to obtain a decomposition of the form A = QRj, where Q is an m x r matrix with
6. Geometrical Considerations 25 orthonormal columns and Rj is anrxfc submatrix whose rows are the r nonnull rows of a k x k upper triangular matrix R having r positive diagonal elements and k — r null rows. Solution. Denote the first /rth columns of A by aj a*, respectively. Then, according to the results of Exercise 4, there exist scalars x,y (i < j = \ k) such that the k column vectors bi b* defined recursively by the equalities bi=aj, b2 = a2-*i2bj, by = ay - xj-\jbj-\ *iybi, b* = a* — **-i,jfcb*_i xutbi, or equivalently by the equalities ai =bi 32 = b2+^12bi, ay = by + Xj-ijbj-i + • • • + jriybi, a* = b* + xt-i,kbk-i + • • • + *i*bi, form an orthogonal set. Further, r of the vectors bi b*. say the si th srth of them, are nonnull, and, for j = 1 k and for those i < j such that b,- is nonnull, jc/y is given uniquely by _ ayb/ Xij-^bi' Now, let B represent the m x k matrix whose first,..., /:th columns are bj b*, respectively, and let X represent the k x k unit upper triangular matrix whose ij th element is (for i < j = 1 k) Jt/y. Then, observing that the first column of BX is bj andthat(fory = 2,..., k)thejthcolumnofBXisby+xy_i,yby_i-H • -+*iybi and recalling result (2.2.9), we find that A = BX = B!Xi, where Bi is the m x r submatrix (of B) whose columns are the sith srth columns of B and X\ is the r x k submatrix (of X) whose rows are the sith, ..., srthrows of X. And, the decomposition A = BiXi can be reexpressed as A = QRi,
26 6. Geometrical Considerations where Q = BjD, with D = diag(|| bSl ||_1 || bs, H"1), and Kx = EXi, with E = diag(|| bSl II || bSr ||), or equivalently where Q is the m x r matrix with ;th column || bSj ||_1 bs. and Ri = {/,;} is the r x k matrix with {lib*, II**,;. for; > 5/, II bf|.||, for7=^-, 0, for j < si. Moreover, the columns of Q are orthonormal, and Rj is an r x k submatrix whose rows are the r nonnull rows of a k x k upper triangular matrix R having r positive diagonal elements and n — r null rows — the s\ th,..., srth rows of R (which are the nonnull rows) are respectively the first,..., rth rows of Ri.
7 Linear Systems: Consistency and Compatibility EXERCISE 1. (a) Let A represent an in x n matrix, C an n x q matrix, and B a q x p matrix. Show that if rank(AC) = rank(C), then ft(ACB) = ft(CB) and rank(ACB) = rank(CB) and that if rank(CB) = rank(C), then C(ACB) = C(AC) and rank(ACB) = rank(AC). (b) Let A and B represent mxn matrices. (1) Show that if C is an r x q matrix and Da<y x/H matrix such that rank(CD) = rank(D),thenCDA = CDB implies DA = DB. [Hint. To show that DA = DB, it suffices to show that rank[D(A - B)] = 0.} (2) Similarly, show that if C is an n x q matrix and Da^xp matrix such that rank(CD) = rank(C). then ACD = BCD implies AC = BC. Solution, (a) It is clear from Corollary 4.2.3 that ft(ACB) c ft(CB) and C(ACB) C C(AC). Now, suppose that rank(AC) = rank(C). Then, according to Corollary 4.4.7, 7£(AC) = 71(C), and it follows from Lemma 4.2.2 that C = LAC for some matrix L. Thus, 7£(CB) = Te(LACB) C 7£(ACB), implying that ft(ACB) = ft(CB) [which implies, in turn, that rank(ACB) = rank(CB)]. Similarly, if rank(CB) = rank(C), then C(CB) = C(C), in which case C = CBR for some matrix R, implying that C(AC) = C(ACBR) c C(ACB) and hence that C(ACB) = C(AC) [and rank(ACB) = rank(AC)].
28 7. Linear Systems: Consistency and Compatibility (b) Let F = A - B. (1) Suppose that rank(CD) = rank(D). Then, if CDA = CDB, we find, in light of Part (a), that rank(DF) = rank(CDF) = rank(CDA - CDB) = rank(O) = 0, implying that DF = 0 or equivalently that DA = DB. (2) Similarly, suppose that rank(CD) = rank(C). Then, if ACD = BCD, we find, in light of Part (a), that rank(FC) = rank(FCD) = rank(ACD - BCD) = rank(0) = 0, implying that FC = 0 or equivalently that AC = BC.
8 Inverse Matrices EXERCISE 1. Let A represent anmxn matrix. Show that (a) if A has a right inverse, then n > m and (b) if A has a left inverse, then m>n. Solution, (a) If A has a right inverse, then, according to Lemma 8.1.1, rank(A) = m, and, since (according to Lemma 4.4.3) n > rank(A), it follows that n>m. (b) Similarly, if A has a left inverse, then according to Lemma 8.1.1, rank(A) = /?, and, since (according to Lemma 4.4.3) m > rank(A), it follows that m > n. EXERCISE 2. Annxn matrix A is said to be involutory if A2 = I, that is, if A is invertible and is its own inverse. (a) Show that an n x n matrix A is involutory if and only if (I — A) (I + A) = 0. (b) Show that a 2 x 2 matrix A = I , I is involutory if and only if (1) a2 + be = 1 and d = -a or (2) b = c = 0 and d = a = ±1. Solution, (a) Clearly, (I-A)(I + A) = I-A + (I-A)A = I-A + A-A2 = I-A2. Thus, (I-A)(I + A) = 0 & I-A2 = 0 <& A2 = I. (b) Clearly, 2 _ (a2 + bc ab + bd\ ~\ac + cd bc + d2) '
30 8. Inverse Matrices And, if Condition (1) or (2) is satisfied, it is easy to see that A is involutory. Conversely, suppose that A is involutory. Then, ab = —db and ac — —dc, implying that d = — a or b = c = 0. Moreover, a2 + be = 1 and d2 +bc = 1. Consequently, if d = —a. Condition (1) is satisfied. Alternatively, if b = c = 0, then dr = a2 = 1, implying that d = a = ±1 (in which case Condition (2) is satisfied) or that d = — a = ±1 (in which case Condition (1) is satisfied). EXERCISE 3. Let A represent an«xn nonnull symmetric matrix, and let B represent an n x r matrix of full column rank r and Tanrxw matrix of full row rank r such that A = BT. Show that the r x r matrix TB is nonsingular. (Hint. Observe that A'A = A2 = BTBT.) Solution. Since A'A = A2 = BTBT, we have (in light of Corollaries 7.4.5 and 8.3.4) that rank(BTBT) = rank(A'A) = rank(A) = r. And, making use of Lemma 8.3.2, we find that rank(BTBT) = rank(TBT) = rank(TB). Thus, rank(TB) = r. EXERCISE 4. Let A represent an n x n matrix, and partition A as A = (Aj, A2). (a) Show that if A is invertible and A-1 is partitioned as A-1 = I -J J (where Bj has the same number of rows as A1 has columns), then BjAi=I, BjA2 = 0. B2Aj=0, B2A2 = I, (E.l) AjBj = I - A2B2, A2B2 = I - AiBj . (E.2) (b) Show that if A is orthogonal, then a; Aj = i, a;a2 = 0, a;aj = o a2a2 = i, (E.3) Ai A', = I - A2A2. A2A2 = I - Aj A', (E.4) Solution, (a) To establish results (E.l) and (E.2). it suffices to observe that if A is invertible and A-1 is partitioned as A-1 = I -J ], then and AjB, +A2B2 = (A,,A2)(g^ = AA-1=I.
8. Inverse Matrices 31 (b) Results (E.3) and (E.4) can be obtained as a special case of results (E.1) and (E.2) by observing that if A is orthogonal, then A is invertible and A-1 = A' = (¾) EXERCISE 5. Let A represent an m x n nonnull matrix of rank r. Show that there exists an m x m orthogonal matrix whose first /• columns span C(A). Solution. According to Theorem 6.4.3, there exist r m-dimensional vectors that are orthonormal with respect to the usual inner product for 1Zmxl and form a basis for C(A). And, according to Theorem 6.4.5, there exist m — r additional m-dimensional vectors, say br+i b,„, such that bj br, br+j b,„ are orthonormal with respect to the usual inner product for 1Zm x l and form a basis for 7£",xl. Clearly, the m x m matrix whose first rth, (r+ l)th mth columns are respectively bj b,, br+j b,„ is orthogonal, and its first r columns span C(A). EXERCISE 6. Let T represent znn x n triangular matrix. Show that rank(T) is greater than or equal to the number of nonzero diagonal elements in T. Solution. Suppose that T has w nonzero diagonal elements and that they are located in the i\ th, iSth /w,th rows of T. Let T* represent the m x m submatrix obtained by striking out all of the rows and columns of T except the «i th, /oth, ..., imth rows and columns. Then, T* is triangular, and the diagonal elements of T+, which are identical to the nth, /2th /,„th diagonal elements of T, are all nonzero. Thus, it follows from Corollary 8.5.6 that rank(T*) = /«. We conclude, on the basis of Theorem 4.4.10, that rank(T) > m. EXERCISE 7. Let A = /An A,2 0 A22 Vo 0 B = /Bji 0 ... 0 \ B2i B22 0 \Bri Br2 Brr/ represent respectively an n x n upper block-triangular matrix whose //th block A/y is of dimensions n\ xiij(j >i = 1 r) and an n x n lower block-triangular matrix whose ijth block B/y is of dimensions «,• xnj (j <i = 1 r). (a) Assume that A and B are invertible, and "recall" that A"l = /Fn F12 0 F22 \0 0 F2r B-' = /Gn G21 0 G22 0 \ 0 \Grj Gr2 ••• Gri/
32 8. Inverse Matrices where J VH = AJ}1, F(/ = -A::1 £ \ik¥kj (j > i = 1,..., r), (*) GK = B^1, G/y = -¾1 ^B«Gjy 0' < i = 1,..., r). (**) k=j Show that the submatrices F,y (j > i = 1,..., r) and G,y 0" < i = 1 r) are also expressible as 7-1 Fyy = A7\ Fy = -(^FrtAjy)A^ (i < y = 1...., r), (E.5) *¥ Gyy = BTy', G/y = -< ]T G/itB^Bj/ (« > y = 1,..., r). (E.6) k=j+l Do so by applying results (**) and (*) to A' and B\ respectively. (b) Formulas (*) form the basis for an algorithm for computing A~l in r steps: the first step is to compute the matrix Frr = A"1; the (r — i 4- 1 )th step is to compute the matrices F,/, F/.,-+i F,> from formulas (*) (i = r - 1, r — 2 1). Similarly, formulas (**) form the basis for an algorithm for computing B_1 in r steps: the first step is to compute the matrix Gj i = BJ",1; the i th step is to compute the matrices Gn, G/2,..., G,,- from formulas (**) (i — 2,..., r). Describe how formulas (E.5) and (E.6) in Part (a) can be used to devise r-step algorithms for computing A~l and B-1, and indicate how these algorithms differ from those based on formulas (*) and (**). Solution, (a) Clearly, it suffices to show that (A')"1 = where /F'u 0 ... 0 \ 9 Foo 0 \K 1¾ - Kr) , (B') /1-1 = /Gii G21 0 G22 \0 0 GU G'J y-i fjj^lA'jjr1. ^. = -^,.)^(¾¾) d<j = l r), G}y = (B};rl, 6^ = -(¾)^^¾¾) (/>y = l r). k=j+\ or equivalently (after relabeling the i and j subscripts) where F;7 = (A;/r1, 1^ = -^,.)-1^¾) (y</ = i d,
8. Inverse Matrices 33 Jt=j+i Upon observing that A' = (M)X 0 12 22 \K a;, o \ 0 KrJ B' = /B'll B2I ••• Brl\ o b„ ... b' \o KrJ the validity of these formulas for (A') x and (B') l is seen to be an immediate consequence of formulas (**) and (*), respectively. (b) To compute A-1, we can employ an r-step algorithm, whose first step is to compute Fj | = A^1 and whose yth step is to compute the matrices F|y-, F2/ Fjj from formulas (E.5) {j = 2 r). To compute B_1, we can employ an r-step algorithm, whose first step is to compute Grr = B"1 and whose (r — j + 1 )th step is to compute the matrices Gyy, Gy+i.y Grj from formulas (E.6) (j = r — 1. /• — 2 1). These algorithms differ from those based on formulas (*) and (**) in that they generate A-1 and B_1 one "column" of blocks at a time, rather than one "row" at a time.
9 Generalized Inverses EXERCISE 1. Let A represent any m x n matrix and B any m x p matrix. Show that if AHB = B for some n x m matrix H, then AGB = B for every generalized inverse G of A. Solution. Suppose that AHB = B for some n x m matrix H, and let G represent an arbitrary generalized inverse of A. Then, AGB = AGAHB = AHB = B. [Or, alternatively, observe that HB is a solution to the linear system AX = B (in X), so that this linear system is consistent and it follows from Theorem 9.1.2 that GB is a solution to AX = B or equivalently that AGB = B.j EXERCISE 2. (a) Let A represent an m x n matrix. Show that any n x m matrix X such that A'AX = A' is a generalized inverse of A and similarly that any n x m matrix Y such that AA'Y' = A is a generalized inverse of A. (b) Use Part (a), together with the result that (for any matrix A) the linear system A'AX = A' (in X) is consistent, to conclude that every matrix has at least one generalized inverse. Solution, (a) Suppose that X is such that A'AX = A'. Then, A'AXA = A'A = A'AI, and it follows from Corollary 5.3.3 that AXA = AI = A
36 9. Generalized Inverses (i.e., that X is a generalized inverse of A). Similarly, if Y is such that AA'Y' = A, then AA'Y'A' = AA' = AA'I, implying that A'Y'A' = A'l = A' and hence that AYA = (A'Y'A')' = (A')' = A. (b) The consistency (for any matrix A) of the linear system A'AX = A' implies that corresponding to any matrix A, there exists a matrix X such that A'AX = A' (and a matrix Y such that AA'Y' = A). Thus, it follows from Part (a) that every matrix has at least one generalized inverse. EXERCISE 3. Let A represent an m x n nonnull matrix, let B represent a matrix of full column rank and T a matrix of full row rank such that A = BT, and let L represent a left inverse of B and R a right inverse of T. (a) Show that the matrix R(B'B)-1 R' is a generalized inverse of the matrix A'A and that the matrix L'Cn")-1!, is a generalized inverse of the matrix AA'. (b) Show that if A is symmetric, then the matrix R(TB)_1L is a generalized inverse of the matrix A2. (If A is symmetric, then it follows from the result of Exercise 8.3 that TB is nonsingular.) Solution, (a) Clearly, A'AtRfB'Br'R'JA'A = T,B,BTR(B,B)~1R,rB,BT = TB'BKB'Br'CTR/B'BT = T'l'B'BT = T'B'BT = A'A, and similarly AA'tlATTr^lAA' = BTT'B'L'(TT'r1LBTT'B' = BTr(LB)'(TT'r1ITT'B' = BTT'I'B' = BTT'B' = AA'. (b) Clearly, A2[R(TB)_1LJA2 = BTBTR(TB)_1LBTBT = BTBI(TB)_1ITBT = BTBT = A2. EXERCISE 4. A generalized inverse, say G, of an m x n matrix A of rank r can be obtained by an approach consisting of (1) finding r linearly independent rows (of A), say rows i\*h ir (where i\ < /2 < • • • < ir). and r linearly independent columns, say j\, 72 jr (where j\ < jz < ■ ■ ■ < jr), (2) inverting the submatrix, say Bj j, of A obtained by striking out all of the rows and columns
9. Generalized Inverses 37 (of A) save rows i\J2 'V and columns ji,j2 jr. and (3) taking (for ^=1,2 r and t = 1,2 /•) the jsi,th element of G to be the .mh element of BJ'j1 and taking its other [n - /*)(/« - r) elements to be 0. Use this approach to find a generalized inverse of the matrix A = /0 0 0 0 o\ 2 -1 3 Solution. The second and third columns of A are linearly independent (as can be easily verified), implying (since the first column is null) that r = 2. Choose, for example, the linearly independent rows and linearly independent columns so that i"i = 2 and ii = 4 — clearly, the second and fourth rows of A are linearly independent— and j\ = 2 and j2 = 3. Then, --G ')■ Applying formula (8.1.2) for the inverse of a 2 x 2 nonsingular matrix, we find that 87/ = (1/6)(4 1)- Thus, one generalized inverse of A is G = (1/6) /0 0 0 0 0\ (0 3 0-2 0 . \0 -3 0 4 0/ EXERCISE 5. Let A represent aninxn nonnull matrix of rank r. Take B and K to be nonsingular matrices (of orders m and /2, respectively) such that -•(i !)' (the existence of which is guaranteed). Show (a) that an n x m matrix G is a generalized inverse of A if and only if G is expressible in the form -My »)"■' (E.1) for some r x (m - r) matrix U, (n -r)xr matrix V, and (n -r)x (m - r) matrix W, and (b) that distinct choices for U, V, and/or W lead to distinct generalized inverses.
38 9. Generalized Inverses Solution, (a) Let H = KGB, and partition H as h=(h2: H,2\ H22>r where Hn is of dimensions r x r. Clearly, G is a generalized inverse of A if and only if •ft 5—(i >=*(' >■ or equivalently (since B and K are nonsingular) if and only if ft ;)»ft j)-ft ;)• and hence if and only if Hn = I. Moreover, if G is expressible in the form (E.1), then H = IS w J, so that Hn = I. Conversely, if Hn = I, then G-«-«r'-Er>(i S)b-=k-(^ ;)r-. with U = H12, V = H21, and W = H22. so that G is expressible in the form (E. 1). We conclude that G is a generalized inverse of A if and only if G is expressible in the form (E.1). (b) Let G, =1^(^ J^B-'andGa^K-1^ JJQb"K where Ui and U2 are r x (m — r) matrices, Vi and V2 are (/? — r) x r matrices, and Wi and W2 are (n -r)x {m - r) matrices. Then, Gi = G2 only if KGjB = KG2B, or equivalently only if (lr V\\_(lr U2\ \yi w,;-^v2 w2;* that is, only if U2 = Ui, V2 = Vi, and W2 = W|. EXERCISE 6. Let k represent a nonzero scalar. For any matrix A, (l/k)\~ is a generalized inverse of the matrix A-A. Generalize this result to partitioned matrices of the form (A, AB) and ( ,.r ), where A is an in x n matrix, B an m x p matrix, and C a q x n matrix. Do so by showing (1) that, for any generalized inverse \r] of the partitioned matrix (A, B) (where Gn is of dimensions n x /»), I _,' J is a generalized inverse of (A, kh) and (2) that, for any generalized inverse (Hj, H2)
9. Generalized Inverses 39 of the partitioned matrix (cj (where Hi is of dimensions n x m) (Hi,*-1H2) is a generalized inverse of (. r ). Solution. (1) Clearly, (A.a»-(4.B)(S i)- Thus, it follows from Part (2) of Lemma 9.2.4 that the matrix (h 0 \"' /G,\ _ /I„ 0 \ /G,\ _ / G, \ \o a,) ^"l« t-'ijVftJ "Vt-'Gj; is a generalized inverse of (A, kB). (2) Similarly, (*c) = (om aj(c)' Thus, it follows from Part (1) of Lemma 9.2.4 that the matrix ""••^ft" i)",=(H"H2)(^ A) = (H"ft"'H2) is a generalized inverse of I , p J. EXERCISE 7. Let T represent anmxp matrix and W an /2 x q matrix. (a) Show that, unless T and W are both nonsingular, there exist generalized (T 0\ /r 0 \ ft w J that are not of the form I ft w_ J. [Hint. Make use of the result that, for any m x n matrix A and for any particular generalized inverse G of A, an n xm matrix G* is a generalized inverse of A if and only if G* = G + (I - GA)T + S(I - AG) for some n x m matrices T and S.] (b) Take U to be anmxq matrix and V an n x p matrix such that C(U) c C(T) and TZ(V) c ft(T), define Q = W - VT~U, and "recall" that the partitioned matrix /T-+T-UQ-VT- -T-UQ-\ V -Q-VT- Q- ) {*} (T U\ J. Generalize the result of Part (a) by showing that, unless T and Q are both nonsingular, there exist generalized inverses of (J, w) that are not of the form (*). [Hint. Use Part (a), together with the result that, for any r x s matrix B, any r xr nonsingular matrix A, and any s x s
40 9. Generalized Inverses nonsingular matrix C, a matrix G is a generalized inverse of ABC if and only if G = C^HA-1 for some generalized inverse H of B.] Solution, (a) Making use of the result cited in the hint, we find that, for any p x n matrix X and q x m matrix Y, the partitioned matrix , 0 W + VY oj[(o iJ~(o w)(o w-jj _/ T~ (IP-T-T)X\ -Vy(Ihi-tt-) w- ) /T 0\ is a generalized inverse of I ft w J. If T is not nonsingular, then either I—T T ^ 0 or I - TT_ ^ 0, as is evident from Corollary 8.1.2. Moreover, if I - T~T # 0, then X can be chosen so that (I - T~T) X # 0, and similarly if I - TT_ # 0, then Y can be chosen so that Y(I - TT~) # 0. We conclude that if T is not (T 0\ ft w J that is not of (T~ 0 \ ft w- )' ^ ^°^ows ^0111 an analogous argument that if W is not (T 0\ ft w J that is not of the form (TQ ^_J. (b) Suppose that either T or Q is not nonsingular, and assume (for purposes of (T U\ v w) *s °^ l^e form (*). Upon observing that (in light of Lemma 9.3.5) /T 0\_/ I 0\ /T u\ /I -T~m vo q) ~ v-vT- i) \\ wj vo i J (I 0\ /I — T~U\ VT_ T I and I ft _ J are nonsingular and upon applying the result cited in the hint and making use of (T 0\ ft n J is of the form /I -T-U\-1 /T-+T-UQVT- -T~UQ-\/ I 0\_1 V0 I ) \ -Q-VT- Q- J^-VT- l) _/I T~U\ /T +TUQ-VT- -T~UQ-\/ I 0\ \0 I A -Q-VT" Q" J^VT" l) -(V <?-)■
9. Generalized Inverses 41 which contradicts Part (a). We conclude that, unless T and Q are both nonsingular, (T U\ v w I that are not of the form (*). EXERCISE 8. Let T represent an m x p matrix, U an m x q matrix, V an n x p (7 U\ v w J, and define Q = W-VT~U. (a) Show that the matrix , /T- + T- UQ-VT- -T-UQ- VT- Q~ (*) is a generalized inverse of the matrix A if and only if (1) (I-TT-)U(I-Q-Q) = 0, (2) (I-QQ-)V(I-T-T) = 0, and (3) (I - TT-)UQ-V(I - T~T) = 0. (b) Verify that (together) the two conditions C(V) C C(7) and ft(V) C 11(7) imply Conditions (1) - (3) of Part (a). (c) Exhibit matrices T, U, V, and W that (regardless of how the generalized inverses T~ and Q~ are chosen) satisfy Conditions (1) - (3) of Part (a) but do not satisfy (both of) the conditions C(V) C C(7) and 1Z(\) C 11(7). Solution, (a) It is a straightforward exercise to show that /T + (I - TT-)UQ-V(I - T-T) U - (I - TT-)U(I - Q~Q)\ A° \ V-d-QQ-)V(I-T-T) W )' Thus, AGA = A if and only if Conditions (1)-(3) are satisfied. (b) Suppose that C(U) C C(T)andft(V) c 11(7). Then, it follows from Lemma 9.3.5 that (I - TT~ )U = 0 and V(I - T~T) = 0, and hence that Conditions (1) -(3) are satisfied. (c) Take T = 0 and U = 0, take W to be an arbitrary nonnull matrix, and take V to be any nonnull matrix such thatC(V) C C(W). Conditions (1) and (3) are clearly satisfied. Moreover, Q = W, and (in light of Lemma 9.3.5) (I - QQ~)V = 0, so that condition (2) is also satisfied. On the other hand, the condition 7£(V) C 11(7) is obviously not satisfied. EXERCISE 9. Suppose that a matrix A is partitioned as (An An Ai3\ A21 A22 A23 I A31 A32 A33/ and that C(Ai2) C C(k\\) and ft(A2i) C ft(An). Take Q to be the Schur
42 9. Generalized Inverses complement of An in A relative to An, and partition Q as g VQ21 <w (where Qn, Q12, Q21. and Q22 are of the same dimensions as A22, A23, A32, and A33, respectively), so that Qn = A22 - A2iAJ"j A12, Q12 = A23 - A2iAJ"jAi3, Q21 = A32 — A3iAj"j A12, and Q22 = A33 — AsiA^A^. Let \ -QnA2iAri Q?i / ' Define T = ^1 £j*\ U = (£*3Y and V = (A31, A32), or equivalently define T, U, and V to satisfy Show that (1) G is a generalized inverse of T; (2) the Schur complement Q22 - Q2iQi\Qi2 of Qn in Q relative to Q^ equals the Schur complement A33 — VGU of T in A relative to G; and (3) pit - fAnAi3 ~ AuA'2QuQi2>\ w-{ QHQiz r VG = (A3|A71-Q2iQ7IA2iA7lf CbQFi)- (b) Let A represent an 1? x n matrix (where n > 2), let n\ /1* represent positive integers such that n \ -\ h«A- = " (where k > 2), and (for / = 1 k) let /?* = «i -\ H /if. Define (for 1 = 1 k) A,- to be the leading principal submatrixof A of order/?* anddefine(for/ = 1 k— 1)U,- to be the n*x (/?-//*) matrix obtained by striking out all of the rows and columns of A except the first n* rows and the last n — n* columns, V, to be the (n — /?*) x n* matrix obtained by striking out all of the rows and columns of A except the last n — n* rows and first n* columns, and W/ to be the (// — //*) x (n — //*) submatnx obtained by striking out all of the rows and columns of A except the last n — n* rows and columns, so that (for / = 1 k-\) Suppose that (for / = 1 k- 1)C(U,-) CC(A,-) and7?(V,-) C 'E(A,-). Let R(i) _ /Bn Bi'A
9. Generalized Inverses 43 <' = 1 * - 1) and B<*> = Eft. where B\\} = A", B^ = A~U,, B<i> = ViAf. andB^} = Wi-ViAfU, and where (for/ > 2)6^,6^, B^, and B^ are defined recursively by partitioning B^-1*, B^,-0, and B^-0 as B?2-,) = (X{'-I),X«'-,,)i bw-d./y!'-1^ B(i-i,_/Q(I'r1) Q(,rn\ 21 "^-'V' B22 -Wr0 <£l7 (in such a way that X(,'_1) has n,- columns, Y}1'"" has m rows, and Q'/j-" is of dimensions n,- x «,) and (using Q7i"~" to represent a generalized inverse of Q',',-") by taking 11 "I -vx-1* or."-0 j" Show that (1) Bj/ is a generalized inverse of A,- (/ = 1 fc); (2) B^ is the Schur complement of A,- in A relative to Bf/ (/ = 1 A:—1);(3) B^ = BJ'/U/ and Bi'/ = V/fift (/ = 1 k-l). [Note. The recursive formulas given in Part (b) for the sequence of matrices B(1\ ..., B(*_1\ Blk) can be used to generate B(*-1) in k — 1 steps or to generate Ba) in k steps — the formula for generating B(,) from B(,_1) involves a generalized inverse of the n,- x w,- matrix Q(|'|_l). The various parts of B(*-,) consist of a generalized inverse B(|*-I) of A*_j, the Schur complement B22_1) of A*_i in A relative to B|*-1\ a solution B(,2~ } of the linear system A^-jX = U*_i (in X), and a solution B2*~n of the linear system YA*_i = V*_j (in Y). The matrix B{k) is a generalized inverse of A. In the special case where w,- = 1, the process of generating the elements of the n x n matrix B(,) from those of the n x n matrix Bl,_1) is called a sweep operation — see, e.g.. Goodnight (1979).] Solution, (a) (1) That G is a generalized inverse of T is evident upon setting T = An,U = A12, V = A2!,andW = A22 in formula (6.2a) of Theorem 9.6.1 [or equivalently in formula (*) of Exercise 7 or 8]—the conditions C(Aj2) C C(An) and 7£(A2i) C 7£(An) insure that this formula is applicable. (2) Q22-Q2iQnQi2
44 9. Generalized Inverses = A33 - AaiA^Au - (A32 - A3iA„Ai2)Q71(A23 - A2iAJ",Ai3) = A33 - A3i(A7, + Aj"1Ai2Q71A2iA71)Ai3 -A3i(-AJ"1Ai2Q71)A23 - A32(-Q71A2iA71)Ai3 - A32Q7,A23 = A33-VGU. (3) Partition GU as GU = (*l\ and VG as VG = (Y,, Y2) (where Xu X2, Yi, and Y2 are of the same dimensions as A13, A23, A31, and A32, respectively). Then, Xi = (Af, + A" AizQ^AziAf^Ais + (-A-jA^Q^Aza = Aj'jA^ - AJ"1Ai2QJ"1(A23 - A2iAJ",Ai3) = A^Ai3 - A^A^Q^Qu and X2 = (-Q7, A21A-)A,3 + Q^Azs = Qr,(A23 - A21A-A,3) = 07^,2. It can be established in similar fashion that Yi = A31A^ — (^iQ^^iA^, and Y2 = Q2iQn (b) The proof of results (1), (2), and (3) is by mathematical induction. By definition, B(jj is a generalized inverse of Ai, B^ is the Schur complement of Ai in A relative to B^, and BJjf = B^Ui and B^ = ViBff. Suppose now that B^-1* is a generalized inverse of A,-_i, that B^-1* is the Schur complement of A,_i in A relative to Bl^l\ and that B^-1* = BJ'f^H-i and B^,-0 = V.-ififf0 (where 2 < i < k - 1). Partition A/, U/, and V,- as (where AJ3 has n*_x rows and A^-1* has n*_Y columns). Then, clearly, U^^A^) and V/_, = teV so that X?-" = Bj'fX"". X2_1) = B?f "Afc-" Y«~l) = AjfXf "• and Y2/_,) = A3/~1)B(1/~,). Thus, it follows from Part (a) that B^ is a generalized inverse of A/, that B^ is the Schur complement of A/ in A relative to BJp and that B55 = B(//U/ and B^ = V/B^. We conclude (based on mathematical induction) that (for / = 1 k—l) B^ is a generalized inverse of A/, B^ is the Schur complement of A,- in A relative to
9. Generalized Inverses 45 B^andB^ = Bj'/U, andB^/ = V/B*//. Moreover, sinceB(t*~n is a generalized inverse of A*_,, since Qf~l) = Bj*"" and Bi2~~n is the Schur complement of A*_, in A relative to B1*"". since x]*"0 = hf2~l) = B^^Ujt-, and Yj*"" = B21 = V*-iB(n~ \ and since A* = A, it is evident upon setting T = A*_i, U = U*_i, V = V*_i, and W = Wjt_i in formula (6.2a) of Theorem 9.6.1 [or equivalently in formula (*) of Exercise (7) or (8)] that B^ is a generalized inverse of A*. EXERCISE 10. Let T represent an m x p matrix and W ann x q matrix, and let G = I p l r j (where G\ \ is of dimensions pxm) represent an arbitrary (T 0\ ft W/* Show that Gn is a generalized inverse of T and G22 a generalized inverse of W. Show also that TG12W = 0 and WG21T = 0. Solution. Clearly, ft 0>i-A-AGA-|'TG,, T^A-f70'17 TG,2W "i \o vj) " A ~ AKyA ~ \WG21 WG22/ VWG21T wg22w;* Thus, TGnT = T (i.e., Gn is a generalized inverse of T), WG22W = W (i.e., G22 is a generalized inverse of W), TG12W = 0, and WG21T = 0. EXERCISE 11. Let T represent an m x p matrix, U an m x q matrix, V an n x p matrix, and W an n x q matrix, and define Q = W — VT~U. Suppose that C(U) C C(T) and TZ(V) C ft(T). Prove that for any generalized inverse G= lru r 12 ) of the partitioned matrix (v w I, the (q x n) submatrix G22 is a generalized inverse of Q. Do so via an approach that consists of showing that / I 0\/T U\/I -T"U\_/T 0\ V-vr i)\v w;\o i )~ \o q) and of then using the result cited in the hint for Part (b) of Exercise 7, along with the result of Exercise 10. Solution. Observing (in light of Lemma 9.3.5) that V - VT~T = 0 and that U - TT_U = 0, we find that / I 0\/T U\/I -T-U\/T U\/I -T-U\ ^_VT- IJVV V/)\0 I )-\0 Q^VO I ) Moreover, recalling Lemma 8.5.2, it follows from the result cited in the hint for Part (b) of Exercise 7, or (equivalently) from Part (3) of Lemma 9.2.4, that the
46 9. Generalized Inverses matrix /I -T-UV'/Gn G,2\/ I OV1 ^0 I ) ^G2i G22A-VT~ V /I T-U\/G,i Gi2\/ I 0\ -^o i ;vg21 G22Avt~ V -( Gu + T-UG2i + G,2VT- + T-UG22VT- G12 + T-UGzA G21+G22VT- G22 ) (T 0\ ft n J. Based on the result of Exercise 10, we conclude that G22 is a generalized inverse of Q.. EXERCISE 12. Let T represent an m x p matrix, U an m x q matrix, V an n x p /T U\ matrix, and W an n xq matrix, and take A = I v w I. (a) Define Q = W - VT~U, and let G = Qjjj1 ^2Y where Gj 1 = T" + T-UQ-VT-, Gi2 = -T-UQ-, G2i = -Q-VT", and G22 = Q~ Show that the matrix Gu -Gi2Gj2G2i is a generalized inverse of T. {Note. If the conditions C(U) C C(T) and 7l(V) C 7£(T) are satisfied or more generally if Conditions (1) - (3) of Part (a) of Exercise 8 are satisfied, then G is a generalized inverse of A}. (b) Show by example that, for some values of T, U, V, and W, there exists a generalized inverse G = I p11 p12 I of A (where Gn is of dimensions p x m, G12 of dimensions p x «, G21 of dimensions q x m, and G22 of dimensions q x n) such that the matrix Gu — G12GJ2G21 is not a generalized inverse of T. Solution, (a) We find that Gn - G,2GJ2G2i = T" + T~UQ-VT- - T~UQ-(Q~rQ-VT" = T" + T~UQ-VT" - T~UQ-VT" = r\ (b) TakeT=H [jY U = 0, V = 0. and W = 0. and take q _ /Gn Gi2\ VG21 G22/
9. Generalized Inverses 47 where Gu = T , G22 = 0, and G12 and G21 are arbitrary. Then, clearly G is a generalized inverse of A. Further, T(Gn - Gi2GJ2G2i)T = T- TG,2G22G2iT, so that Gn — G12GJ2G21 is a generalized inverse of T if and only if TGi2G2~2G2iT = 0. Suppose, for example, that G12» G^, and G21 are chosen so that the (1,1 )th element of G12 G7-> G21 is nonzero (which—since any n x q matrix is a generalized inverse of G22 — is clearly possible). Then, the (1,1 )th element of TGj2G2^G2i T is nonzero and hence TG^GJ-^iT is nonnull. We conclude that Gj 1 -GnGJ;^! is not a generalized inverse of T.
10 Idempotent Matrices EXERCISE 1. Show that if an n x n matrix A is idempotent, then (a) for any n x n nonsingular matrix B, B~l AB is idempotent; and (b) for any integer k greater than or equal to 2, A* = A. Solution. Suppose that A is idempotent. (a) B-'ABCB-'AB^ 8-^8 = 8-^8. (b) The proof is by mathematical induction. By definition, A* = A for k = 2. Suppose that A* = A for k = k*. Then, A**+1=AAr =AA = A, that is, A* = A for k = k* + 1. Thus, for any integer k > 2, A* = A. EXERCISE 2. Let Prepresent anm x n matrix (wherem >n) such that PT = I„, or equivalently an m x n matrix whose columns are orthonormal (with respect to the usual inner product). Show that the m x m symmetric matrix PP7 is idempotent. Solution. Clearly, (PP,)PP/ = PtPW = PlnV = PP7. EXERCISE 3. Show that, for any symmetric idempotent matrix A, the matrix I - 2A is orthogonal. Solution. Clearly, (I-2A)'(I-2A) = (I-2A)(I-2A) = I-2A-2A+4A2 = I-2A-2A+4A = I.
50 10. Idempotent Matrices EXERCISE 4. Let A represent an m x n matrix. Show that if A'A is idempotent, then AA' is idempotent. Solution. Suppose that A'A is idempotent. Then, A'AA'A = A'A = A'AI, implying (in light of Corollary 5.3.3) that AA'A = AI = A and hence that (AA')2 = (AA'A)A' = AA'. EXERCISE 5. Let A represent a symmetric matrix and k an integer greater than or equal to 1. Show that if A*+1 = A*, then A is idempotent. Solution. It suffices to show that, for every integer /// between k and 2 inclusive, Am+l = Aw impHes xm = Am-\ (If A*+l = A* but pj. were nQt equa, tQ A then there would exist an integer m between k and 2 inclusive such that Am+1 = A"1 butA'^A'""1.) Suppose that A"'+1 = A'". Then, since A is symmetric, A'AA'"-1 =A,AA"'-2 (where A0 = I), and it follows from Corollary 5.3.3 that AA"'-1 =AA'W-2 or equivalently that A,M=A,H-1 EXERCISE 6. Let A represent an n x // matrix. Show that (1/2)(1 + A) is idempotent if and only if A is involutory (where involutory is as defined in Exercise 8.2). Solution. Clearly, [(1/2)(1 +A)]2 = (1/4)1+ (1/2)A + (1/4)A2. Thus, (1/2)(1 + A) is idempotent if and only if (1/4)1 + (1/2)A + (1/4)A2 = ( 1/2)1 + (1/2)A, or equivalently if and only if (1/4)A2 = (1/4)1, and hence if and only if A2 = I (i.e., if and only if A is involutory). EXERCISE 7. Let A and B represent // x // symmetric idempotent matrices. Show that if C'(A) = t'(B), then A = B.
10. Idempotent Matrices 51 Solution. Suppose that C(A) = C(B). Then, according to Lemma 4.2.2, A = BR and B = AS for some n xrt matrices R and S. Further, B = B' = (AS)' = S'A' = S'A. Thus, A = BBR = BA = S'AA = S'A = B. EXERCISE 8. Let A represent an r x m matrix and Banmxn matrix. (a) Show that B~A~ is a generalized inverse of AB if and only if A ABB is idempotent. (b) Show that if A has full column rank or B has full row rank, then B~A~ is a generalized inverse of AB. Solution, (a) In light of the definition of a generalized inverse and the definition of an idempotent matrix, it suffices to show that ABB-A~AB = AB if and only if A~ ABB" A" ABB" = A~ABB-. Premultiplication and postmultiplication of both sides of the first of these two equalities by A- and B~, respectively, give the second equality, and premultiplication and postmultiplication of both sides of the second equality by A and B, respectively, give the first equality. Thus, these two equalities are equivalent. (b) Suppose that A has full column rank. Then, according to Lemma 9.2.8, A~ is a left inverse of A (i.e.. A-A = I). It follows that A-ABB- = BB~ and hence, in light of Lemma 10.2.5, that A-ABB- is idempotent. We conclude, on the basis of Part (a), that B~ A~ is a generalized inverse of AB. It follows from an analogous argument that if B has full row rank, then B~ A~ is a generalized inverse of AB. EXERCISE 9. Let T represent an m x p matrix. U an m x q matrix, V an n x p matrix, and W an n x q matrix, and define Q = W — VT~~U. Using the result of Part (a) of Exercise 9.8, together with the result that (for any matrix B) rank(B) = tr (B~B) = tr (BB~), show that if (1) (I-TT-)U(I-CTQ) = 0, (2) (I - QQ~ )V(I - T~T) = 0, and (3) (I - TT" )UQ~ V(I - T~T) = 0, then rank (^ ^ j = rank(T) + rank(Q). (T U\ J, and define G as in Part (a) of Exercise 9.8. Suppose that Conditions (1)-(3) are satisfied. Then, in light of Exercise 9.8, G is
52 10. Idempotent Matrices a generalized inverse of A, and, making use of the result that (for any matrix B) rank(B) = tr (BB~) [which is part of result (2.1)], we find that rank(A) = tr(AG) /TT- - (I - TT-)UQ-VT- (I - TT~)UQ-\ V (I-QQ-)VT- QQ- ) = tr (TT") - tr [(I - TT-JUQ-VT"] + tr (QQ~) = rank(T) - tr [(I - TT~)UQ-VT~] + rank(Q). Moreover, it follows from Condition (3) that (I - TT-)UQ~V = (I - TT-)U(r VT-T and hence that tr [(I - TT-)UQ- VT~] = tr [(1 - TT-)UQ- VTTT-] = tr [TT-(I - TT-)UQ-VT-] = tr [(TT- - TT-)UQ~VT~] = tr(0) = 0. We conclude that rank(A) = rank(T) + rank(Q). EXERCISE 10. Let T represent an m x p matrix, U an m x q matrix, V an n x p /T U\ matrix, and W an« x q matrix, take A = I v w J, and define Q = W- VT U. Further, let Er = I-TT-, F7=I-T-T, X = E7U, Y = VF7, Ey=I-YY~, FX=I-X~X, Z = EyQFx, and Q* = FxZ~Ey . (a) (Meyer 1973, Theorem 3.1) Show that the matrix 0 = ^+02, (E.1) where G,= / T--T-U(I-Q*Q)X-Er -F7Y-(I-QQ*)VT- -FrY(I-QQ*)QXEr FrY(I-QQ*) and V (I-Q*Q)XEr G2 = ^~U^Q*(-VT-, I„),
10. Idempotent Matrices 53 is a generalized inverse of A. (b) (Meyer 1973, Theorem 4.1) Show that rank(A) = rank(T) + rank(X) + rank(Y) + rank(Z). (E.2) [Hint. Use Part (a), together with the result that (for any matrix B) rank(B) = tr(B~B) = tr(BB-).] (c) Show that if C(U) C C(T) and 1Z(\) c K(T), then formula (E.2) for rank(A) reduces to the formula rank(A) = rank(T) + rank(Q), and the formula /T-+T-UQ-VT- -T~UQ-\ ^ -Q-VT- Q" j' (*' which is reexpressible as (T !)+(Tu)Q-'-"-u- can be obtained as a special case of formula (E.1) for a generalized inverse of A. Solution, (a) It can be shown, via some painstaking algebraic manipulation, that AGlA = (v w-qq*q) ^ AG2A=(o qq'q) and hence that AGA = AG!A + AG2A = A. (b) Taking G to be the generalized inverse (E.1), it is easy to show that /TT- + XX-Er| 0 \ Al*-^ oTnltted | YY-+EyQQV' Thus, making use of the result that (for any matrix B) rank(B) = tr (BB~) [which is part of result (2.1)], we find that rank(A) = tr(AG) = tr (TT~) + tr (XX"E7) + tr (YY~) + tr (Ey QQ*) = rank(T) + tr (XX~Er) + rank(Y) + tr (EyQQ*). Moreover, tr(XX-E7) = tr(ErXX~) = tr(ErErUX-) = tr(ErUX-) = tr(XX~) = rank(X),
54 10. Idempotent Matrices and similarly tr(ErQQ*) = trfEyQFxZ'Er) = tr(EYEYQFxZ~) = tr(EyQFxZ-) = tr(ZZ~) = rank(Z). (c) Suppose that C(U) C C(T)andfc(V) C ft(T). Then, it follows from Lemma 9.3.5 that X = 0 and Y = 0. Accordingly, Fx = I and EY = I, implying that Z = Q. Thus, formula (E.2) reduces to rank(A) = rank(T) + rank(Q), [which is formula (9.6.1)]. Clearly, Q* is an arbitrary generalized inverse of Q, and the q x m and p x « null matrices are generalized inverses of X and Y, respectively. Thus, formula (**) can be obtained as a special case of formula (E. 1) by setting X~ = 0 and Y~ = 0 — formula (**) is identical to formula (9.6.2b). and formula (*) identical to formula (9.6.2a) [and to formula (*) of Exercise 9.7 or 9.8].
11 Linear Systems: Solutions EXERCISE 1. Show that, for any matrix A, C(A)=A/"(I-AA-). Solution. Letting x represent a column vector (whose dimension equals the number of rows in A), it follows from Corollary 9.3.6 that x e C(A) if and only if x = AA~x, or equivalently if and only if (I — AA~ )x = 0, and hence if and only if xeAfa- AA~). We conclude that C(A) = N(l - AA~). EXERCISE 2. Show that if Xi X* are solutions to a linear system AX = B (in X) and ci ct are scalars such that J^=i Ci = h mentne matrix J^=i c&i is a solution to AX = B. Solution. If Xi X* are solutions to AX = B and c\ c* are scalars such that Y!l=\ V = L men (k \ k k Y,CiXi ) = 2><AX/) = £>B = B. i=i / i=l i=i EXERCISE 3. Let A and Z represent /2 x n matrices. Suppose that rank(A) = n — 1, and let x and y represent nonnull «-dimensional column vectors such that Ax = 0 and A'y = 0. (a) Show that AZ = 0 if and only if Z = xk' for some ^-dimensional row vector k'.
56 11. Linear Systems: Solutions (b) Show that AZ = ZA = 0 if and only if Z = cxy' for some scalar c. Solution, (a) Suppose that Z = xk' for some row vector k'. Then, AZ = (Ax)k' = 0k; = 0. Conversely, suppose that AZ = 0. Let zy- represent the jth column of Z. Since (in light of Lemma 11.3.1 and Theorem 4.3.9) {x} is abasis forjV(A), zy = kjxfor some scalar kj (/ = 1 //), in which case Z = xk', where k' = (k\,..., k„). (b) Suppose that Z = cxy' for some scalar c. Then, AZ = 0 [as is evident from Part (a)], and ZA = cxy'A = cx(A'y)' = cxO' = 0. Conversely, suppose that AZ = ZA = 0. Then, it follows from Part (a) that Z = xk' for some row vector k'. Moreover, k' = (x'x)-1 (x'x)k' = (x'x)-1x'Z, so that k'A = (x/x)~1x/ZA = 0, implying that A'k = (k'A)' = 0 and hence that k e N{A'). Since (in light of Lemma 11.3.1 and Theorem 4.3.9) {y} is a basis for jV(A;), k = cy for some scalar c. Thus, Z = cxy'. EXERCISE 4. Suppose that AX = B is a nonhomogeneous linear system (in an n x p matrix X). Let s = p[n — rank(A)], and take Zi Zs to be any s nx p matrices that form a basis for the solution space of the homogeneous linear system AZ = 0 (in annx p matrix Z). Define Xo to be any particular solution to AX = B, and let X,- = Xo + Z,- (/ = 1 s). (a) Show that the s + 1 matrices Xo, Xj X5 are linearly independent solutions to AX = B. (b) Show that every solution to AX = B is expressible as a linear combination of Xo, Xj Xs. (c) Show that a linear combination 51/=0 *'X,- of Xo, Xj X5 is a solution to AX = B if and only if the scalars ko,ki ks are such that 51/=0 ^/ = 1- (d) Show that the solution set of AX = B is a proper subset of the linear space sp(Xo,X! Xs). Solution, (a) It follows from Theorem 11.2.3 thatXj X5, likeXo, are solutions toAX = B. For purposes of showing that Xo, Xj X5 are linearly independent, suppose that ko. k\ ks are scalars such that £J=o hX,- = 0. Then, (J> JXo + X>/Zf- = ][>X/ = 0. (S.l) V=0 / /=1 /=1 Consequently, (i>)b=A[(§'ri)Xo+5^]=o>
11. Linear Systems: Solutions 57 implying (since B £ 0) that s 5>«=0 (S.2) /=0 and hence [in light of equality (S.l)] that £?=1 fc/Z/ = 0. Since Zj,..., Zs are linearly independent, we have that k\ = ... = ks = 0, which, together with equality (S.2), further implies that k0 = 0. We conclude that X0, Xj Xs are linearly independent. (b) Let X* represent any solution to AX = B. Then, according to Theorem 11.2.3. X* = X0 + Z* for some solution Z* to AZ = 0. Since Zj Zs form a basis for the solution space of AZ = 0, Z* = £-=1 k\L\ for some scalars *,- AvThus, X* = X0 + £>Z,. = ( 1 - f^ki )Xo + X>x,. /=i \ i=i / i=i (c) We find that Thus, if YlUoh = *• then Ef=o^'X' is a solution to AX = B. Conversely, if 5Z'/=0^'xi is a solution to AX = B, then clearly (£-=0fci)B = B, implying (since B#0) that £f=0*/ = l. (d) It is clear from Part (b) that every solution to AX = B belongs to sp(Xo, Xj, ..., X,). However, not every matrix in sp(Xo, Xj,..., Xs) is a solution to AX = B, as is evident from Part (c). Thus, the solution set of AX = B is a proper subset of sp(Xo, Xj Xs). EXERCISE 5. Suppose that AX = B is a consistent linear system (in an n x p matrix X). Show that if rank(A) < n and rank(B) < p, then there exists a solution X* to AX = B that is not expressible as X* = GB for any generalized inverse G of A. Solution. Suppose that rank(A) < n and rank(B) < p. Then, since the columns of B are linearly dependent, there exists a nonnull vector k] such that Bki = 0, and, according to Theorem 4.3.12, there exist p — 1 /7-dimensional column vectors k2 kp such that the set {kj, k2 kp) is a basis for HP. Define K = (ki, K2), where K2 is the p x (p -1) matrix whose columns are k2 kp. Clearly, the matrix K is nonsingular. Since the columns of A are linearly dependent, there exists a nonnull vector y* such that Ay J = 0. Let Y* = (y*, Y£), where YJ is any solution to the linear system AY2 = BK2 (in Y2). (Since AX = B is consistent, so is AY2 = BK2.) Clearly, AY* = BK.
58 11. Linear Systems: Solutions Define X*=Y*K_1. Then, AX* = AY*K_1 = BKK-1 = B, so that X* is a solution to AX = B. To complete the proof, it suffices to show that X* is not expressible as X* = GB for any generalized inverse G of A. Assume the contrary, that is, assume that X* = GB for some generalized inverse G of A. Then, since (y*, Y|) = Y* = X*K = (X*k,, X*K2), we have that yJ=X*k, =GBk, =0, which (since, by definition, y* is nonnull) establishes a contradiction. EXERCISE 6. Let A represent an m x n matrix and B an m x p matrix. If C is an r x m matrix of full column rank (i.e., of rank m)% then the linear system CAX = CB is equivalent to the linear system AX = B (in X). Use the result of Part (b) of Exercise 7.1 to generalize this result. Solution. If C is an /• x q matrix and D a q x m matrix such that rank(CD) = rank(D), then the linear system CDAX = CDB (in X) is equivalent to the linear system DAX = DB (in X) [as is evident from Part (b) of Exercise 7.1]. EXERCISE 7. Let A represent anmxn matrix, B an hi x p matrix, and C a q x in matrix, and suppose that AX = B and CAX = CB are linear systems (in X). (a) Show that if rank[C(A, B)] = rank(A, B), then CAX = CB is equivalent to AX = B — this result is a generalization of the result that CAX = CB is equivalent to AX = B if C is of full column rank (i.e., of rank m) and also of the result that (for any n x s matrix F. the linear system A'AX = A'AF is equivalent to the linear system AX = AF (in X). (b) Show that if rank[C(A, B)] < rank(A, B) and if CAX = CB is consistent, then the solution set of AX = B is a proper subset of that of CAX = CB (i.e., there exists a solution to CAX = CB that is not a solution to AX = B). (c) Show, by example, that if rank[C(A, B)] < rank(A, B) and if AX = B is inconsistent, then CAX = CB can be either consistent or inconsistent. Solution, (a) Suppose that rank[C(A, B)] = rank(A, B). Then, according to Corollary 4.4.7, ft[C(A, B)] = ft(A, B) and hence ft(A, B) C ft[C(A, B)], implying (in light of Lemma 4.2.2) that (A, B) = LC(A. B) for some matrix L. Therefore, A = LCA and B = LCB. For any solution X* to CAX = CB, we find that AX* = LCAX* = LCB = B.
11. Linear Systems: Solutions 59 Thus, any solution to CAX = CB is a solution to AX = B, and hence (since any solution to AX = B is a solution to CAX = CB) CAX = CB is equivalent to AX = B. (b) Suppose that rank[C(A, B)] < rank(A, B) and that CAX = CB is consistent. And, assume that AX = B is consistent — if AX = B is inconsistent, then clearly the solution set of AX = B is a proper subset of that of CAX = CB. Then, making use of Theorem 7.2.1, we find that rank(A) = rank(A. B) > rank[C(A, B)] = rank(CA, CB) = rank(CA), implying that n - rank(A) < n - rank(CA). (S.3) Let X(> represent any particular solution to AX = B. According to Theorem 11.2.3, the solution set of AX = B is comprised of every n x p matrix X* that is expressible as X*=X0+Z* for some solution Z* to the homogeneous linear system AZ = 0 (in an n x p matrix Z). Similarly, since Xo is also a solution to CAX = CB, the solution set of CAX = CB is comprised of every matrix X* that is expressible as X* = X0+Z* for some solution Z* to the homogeneous linear system CAZ = 0. It follows from Lemma 11.3.2 that the dimension of the solution space of AZ = 0 equals p[n — rank(A)] and the dimension of the solution space of CAZ = 0 equals p[n—rank(CA)]. Clearly, the solution space of AZ = 0 is a subspace of the solution space of CAZ = 0 and hence, in light of inequality (S.3), it is a proper subspace. We conclude that the solution set of AX = B is a proper subset of the solution set ofCAX = CB. (c) Suppose that AX = B is any inconsistent linear system and that C = 0, in which case rank[C(A, B)] = 0 < rank(A, B). Then, CAX = CB is clearly consistent. Alternatively, suppose that a-(?5)..-(9.-c-(..«. in which case AX = B is obviously inconsistent and rank[C(A, B)] = 1 < 2 = rank(A, B). Then, CAX = CB is clearly inconsistent. EXERCISE 8. Let A represent a q x n matrix, B an m x p matrix, and C an m x q matrix; and suppose that the linear system CAX = B (in an n x p matrix X) is
60 11. Linear Systems: Solutions consistent. Show that the value of AX is the same for every solution to CAX = B if and only if rank(CA) = rank(A). Solution. It suffices (in light of Theorem 11.10.1) to show that rank(CA) = rank(A) if and only if ft(A) C ft(CA) or equivalently [since ft(CA C ft(A)] if and only if ft(A) = ft(CA). If 11(A) = ft(CA), then it follows from the very definition of the rank of a matrix that rank(CA) = rank(A). Conversely, if rank(CA) = rank(A), then it follows from Corollary 4.4.7 that ll(\) = ft(CA). EXERCISE 9. Let A represent anmxn matrix, B an m x p matrix, and K an nxq matrix. Verify (1) that if X* and L* are the first and second parts, respectively, of any solution to the linear system (½ :)(9 = 0 (in X and L), then X* is a solution to the linear system AX = B (in X), and L* = K'X*. and, conversely, if X* is any solution to AX = B, then X* and K'X* are the first and second parts, respectively, of some solution to linear system (*); and (2) (restricting attention to the special case where m = n) that If X* and L* are the first and second parts, respectively, of any solution to the linear system /A + KK' -K\/X\ /B\ { -K' iJUJ = W <**> (in X and L), then X* is a solution to the linear system AX = B (in X) and L* = K'X*, and, conversely, if X* is any solution to AX = B, then X* and K'X* are the first and second parts, respectively, of some solution to linear system (**). Solution. (1) Suppose that X* and L* are the first and second parts, respectively, of any solution to linear system (*). Then, clearly -K'X* + L*=0, or equivalently L* = K'X*, and AX*=AX*+0L*=B, so that X* is a solution to AX = B. Conversely, suppose that X* is a solution to AX = B. Then, clearly /A 0\/X*W AX* WB\ \-Kf l) \K'X*J ~ \-K'X* + K'X*; ~\0) ' so that X* and K'X* are the first and second parts, respectively, of some solution to linear system (*). (2) Suppose that X* and L* are the first and second parts, respectively, of any solution to linear system (**). Then, clearly -KX* + L*=0,
11. Linear Systems: Solutions 61 or equivalently L* = K'X*t and AX* = (A + KK')X* - K(K'X*) = (A + KK')X* - KL* = B, so that X* is a solution to AX = B. Conversely, suppose that X* is a solution to AX = B. Then, clearly, /A + KK' -K\ ( X* \ /(A + KK')X* - KK'X*\ _ /AX*\ _ /B\ ^ -K' I j^K'X*;-^ -KT + K'X* )- \ 0 )- \o)' so that X* and K'X* are the first and second parts, respectively, of some solution to linear system (**).
12 Projections and Projection Matrices EXERCISE 1. Let Y represent a matrix in a linear space V, let U and W represent subspaces of V, and take {X\ X^} to be a set of matrices that spans U and {Zi Z,} to be a set that spans W. Verify that Y JL U if and only if Y-X,- = 0 for i = 1 s (i.e., that Y is orthogonal to U if and only if Y is orthogonal to each of the matrices Xi Xs); and, similarly, that U J. W if and only if X,- • Zy = 0 for i = 1 s and j = 1 t (i.e., that U is orthogonal to VV if and only if each of the matrices Xj Xs is orthogonal to each of the matrices Zj Z,). Solution. Suppose that Y _L U. Then, since X,- e1/, we have that Y»X, = 0 (/ = 1 5). Conversely, suppose that Y*X,- = 0 for i = 1 s. For each matrix X eU, there exist scalars c\ cs such that X = c\X\ -\ h csXs, so that Y-X = ci(Y-Xj ) + ••• + cs(Y-X5) = 0. Thus, Y is orthogonal to every matrix in U, that is, Y _L U. The verification of the first assertion is now complete. For purposes of verifying the second assertion, suppose that U -L W. Then, since X,- 6 U and Yy 6 W, we have that X/ -Yy = 0 (/ = 1 s\ j = 1 /). Conversely, suppose that X,- • Zy = 0 for i = 1 5 and j = 1 t. For each matrix XeU, there exist scalars cj cs such that X = cjXj H \-csXs and, for each matrix Z in W, there exist scalars d\ dt such that Z = d\Z\ +
64 12. Projections and Projection Matrices hd,Z,,sothat X-z = J2 JxrY^djz) = £> £>(X,-zy) = o. Thus, U ± W. EXERCISE 2. Let U and V represent subspaces of Tlmxn. Show that if dim(V) > dim(W), then V contains a nonnull matrix that is orthogonal to U. Solution. Let r = dim(t/) and s = dim(V). And, let {Aj,..., Ar) and {Bj B5] represent bases fort/ and V, respectively. Further, define H = [hy] to be the r x s matrix whose //th element equals A,- *By. Now, suppose that s > r. Then, since rank(H) < r < s, there exists ansxl nonnull vector x = [xj } such that Hx = 0. Let C = .V|B| H h x5Bs. Then, C is nonnull. Moreover, for i = 1,..., r, A/-C = A-1(Al-B,) + ...+A,(A/-B,) = ^/i/>Yy. J Since J^j hijXj is the /th element of the vector Hx, £/ htjXj = 0» and hence A/ • C = 0 (/ = 1 r). Thus, C is orthogonal to each of the matrices Aj Ar. We conclude on the basis of Lemma 12.1.1 (or equivalently the result of Exercise 1) that C is orthogonal to U. EXERCISE 3. Let U represent a subspace of the linear space TV" of all /m- dimensional column vectors. Take M to be the subspace of TZmxn defined by We >W if and only ifW = (wj w„) for some vectors wj w„ in t/.Let Z represent the projection (with respect to the usual inner product) of an m x n matrix Y on M, and let X represent any m x p matrix whose columns span U. Show that Z = XB* for any solution B* to the linear system X'XB = X,Y (inB). Solution. Let y,- represent the /th column of Y, and take v/ to be the projection (with respect to the usual inner product) of y,- on U (/ = 1 n). Define V = (V| vH). Then, by definition, (yf- - v,- )'w = 0 for every vector w in U, so that, for every matrix W = (wj w„) in M, » tr[(Y - V)'W] = £(y, - v/Vw,- = 0, /=l implying that Z = V. Now, suppose that B* is a solution to X'XB = X'Y. Then, for / = 1 n, the /th column b* of B* is clearly a solution to the linear system X'Xb,- = X'y,- (in
12. Projections and Projection Matrices 65 b,). We conclude, on the basis of Theorem 12.2.1, that v/ = Xb* (z = 1 n) and hence that Z = V = (v, v„) = (Xbt Xb*) = XB*. EXERCISE 4. The projection (with respect to the usual inner product) of an n-dimensional column vector y on a subspace U of 11" in the special case where n = 3, y = (3, -38/5,74/5)' and U = sp{xi, x2, x3}, with -(3--(1)--(1)- was determined to be the vector (3,22/5,44/5)'—and it was observed that xi and X2 are linearly independent and that X3 = X2 - (l/3)xj, with the consequence that dim(W) = 2. Recompute the projection of y on U (in this special case) by taking X to be the 3 x 2 matrix n and carrying out the following two steps: (1) compute the solution to the normal equations X'Xb = X'y; and (2) postmultiply X by the solution you computed in Step(l). Solution. (1) The normal equations are /45 30\. _/66\ \30 24) \3%)' They have the unique solution ./45 30^/66^/2/15 -l/6\ /66\ _ /37/15\ ^-^30 2V [}*)-\-l/6 1/4) \3ZJ-{-3/2)- (2) The projection of y on U is -(St®*)-®- EXERCISE 5. Let X represent any n x p matrix. If a p x n matrix B* is a solution to the linear system X'XB = X; (in B), then B* is a generalized inverse of X and XB* is symmetric. Show that, conversely, if a p x n matrix G is a generalized inverse of X and if XG is symmetric, then X'XG = X; (i.e., G is a solution to X'XB = X').
66 12. Projections and Projection Matrices Solution. Suppose that G is a generalized inverse of X and XG is symmetric. Then, X'XG = X'(XG)' = (XGX)' = X'. EXERCISE 6. Using the result of Part (b) of Exercise 9.3 (or otherwise), show that, for any nonnull symmetric matrix A, PA = 8(18)^1, where B is any matrix of full column rank and T any matrix of full row rank such that A = BT. (That TB is nonsingular follows from the result of Exercise 8.3.) Solution. Let L represent a left inverse of B and R a right inverse of T. Then, according to Part (b) of Exercise 9.3, the matrix R(TB)-1 L is a generalized inverse of A2 or equivalently (since A is symmetric) of A'A. Thus, PA = AR(TB)_1LA = BTRdBr'LBT = BI(TB)_IIT = B(TB)-1T. EXERCISE 7. Let V represent a /r-dimensional subspace of the linear space 1Z" of all /i-dimensional column vectors. Take X to be any n x p matrix whose columns span V, let U represent a subspace of V, and define A to be the projection matrix for li. Show (1) that a matrix B (of dimensions n x ri) is such that By is the projection of y on li for every y 6 V if and only if B = A + Z* for some solution Z* to the homogeneous linear system X'Z = 0 (in an n x n matrix Z) and (2) that, unless k = /i, there is more than one matrix B such that By is the projection of y on U for every y 6 V. Solution. (1) The vector Ay is the projection of y on U for every y e 1Z". Thus, By is the projection of y on li for every y 6 V if and only if By = Ay for every y e V, or equivalently if and only if BXr = AXr for every p x 1 vector r, and hence (in light of Lemma 2.3.2) if and only if BX = AX. Furthermore, BX = AX if and only if X'(B — A)' = 0, or equivalently if and only if (B — A)' is a solution to the homogeneous linear system X'Z = 0 (in an n x n matrix Z), and hence if and only if B' = A' + Z* for some solution Z* to X'Z = 0, that is, if and only if B = A + Z't for some solution Z* to X'Z = 0. (2) According to Lemma 11.3.2, the solution space of the homogeneous linear system X'Z = 0 (in an n x n matrix Z) is of dimension n[n—rank(X)] = n(n —k). Thus, unless k = /i, there is more than one solution to AZ = 0, and hence [in light of the result of Part (1)] there is more than one matrix B such that By is the projection of y on U for every y e V. EXERCISE 8. Let {A i A*} represent a nonempty linearly independent set of matrices in a linear space V. And, define (as in Gram-Schmidt orthogonalization) k
12. Projections and Projection Matrices nonnull orthogonal linear combinations, say Bj Bj=A,, B2 = A2-.v,2B,, By = Ay - .Yy_I(yBy_, .YjyB,, Bjt = A* - .v*-ijtBjt_i .vu-Bi, where (for i < j = 1 k) AyB, Show that By is the (orthogonal) projection of Ay on some subspace Uj (of V) and describe Uj 0=2 k). Solution. For j = 1 A\ define Cy = || By H_,By (as in Corollary 6.4.2). And, define Wj = sp(Ci Cy). Then, for j = 2 K B,=A,-g^B,. = A;-gf0C,.=Ay-|:(ArC,C, Moreover* upon observing that the set {Ci,..., Cy} is orthonormal and applying result (1.1), we find that ^/^/(Ay •C/jC,- is the projection of Ay on Wy_i. Thus, it follows from Theorem 12.5.8 that By is the projection of Ay on W£_,. And, since (in light of the discussion of Section 6.4b) Wy_j = sp(Ai,..., Ay_j), we conclude that By is the projection of Ay on the orthogonal complement of the subspace (of V) spanned by Aj Ay_i. 67 .., B^, of Ai Ajt as follows:
13 Determinants 1. Let A = f«ll «12 «13 |«14| |«21 | «22 «23 «24 «31 «32 |«33 | «34 ^«41 |«42| «43 «44 (a) Write out all of the pairs that can be formed from the four boxed elements of A. (b) Indicate which of the pairs from Part (a) are positive and which are negative. (c) Use the formula o-rt(l, in ...; w, in) = an{i\, 1;...; in* n) = 0n(z"i,..., in) (in which *i,..., in represents an arbitrary permutation of the first n positive integers) to compute the number of pairs from Part (a) that are negative, and check that the result of this computation is consistent with your answer to Part (b). Solution, (a) and (b) Pair «14, «14, «14, «211 «21, «33, «21 «33 «42 «33 «42 «42 "Sign" + +
70 13. Determinants 4]- (c) 04(4,1,3,2) = 3+0+1 =4 [oralternatively<M2,4,3,1) = 1+2+1 = EXERCISE 2. Consider the n x n matrix A = "Recall" that /a + X X \ * X .v+X X X A" * + |S| = |R| (*) for any n x n matrix R and for any matrix S formed from R by adding to any one of its rows or columns, scalar multiples of one or more other rows or columns; and use this result to show that |A| = X"-j(«a-4-X). (Hint. Add the last n — 1 columns of A to the first column, and then subtract the first row of the resultant matrix from each of the last n — 1 rows). Solution. The matrix obtained from A by adding the last n — 1 columns of A to the first column is B = (nx + X x nx + X x + X \nx + X a- x A+xy The matrix obtained from B by subtracting the first row of B from each of the next i rows is Q = /nx + X .v . 0 X 0 0 HA* + X A' . . X 0 X . A- X 0 0 .. A'+X X 0 . 0 X \nx + X x x + X/ n — 1 — i rows Observing that C/ can be obtained from C/_i by subtracting the first row of C/_i from the (/' +1 )th row and making use of result (*) (or equivalently Theorem 13.2.10) and Lemma 13.1.1, we find that |A| = |B| = |C,| = |C2| = ..- = IC-il = X"-1^- + X).
13. Determinants 71 EXERCISE 3. Let A represent an n x n nonsingular matrix. Show that if the elements of A and A-1 are all integers, then |A| = ±1. Solution. Suppose that the elements of A and A-1 are all integers. Then, it follows from the very definition of a determinant that |A| and |A_I| are both integers. Thus, since (according to Theorem 13.3.7) |A-11 = 1/| A|, |A| and 1/|A| are both integers. We conclude that |A| = ±1. EXERCISE 4. Let T represent an m x m matrix, U an m x n matrix, V an n x m matrix, and W an n x n matrix. Show that if T is nonsingular, then |V W T U U T| W V = (-l)"",|T||W-VT-IU|. Solution. It follows from Theorem 13.2.7 that V Wl T U = (-!)'" T U V W and U T W V = (-!)'" T U V W| Thus, making use of Theorem 13.3.8, we find that = (-1)'""|T||W - VT_1U|. V w T U U T W V EXERCISE 5. Compute the determinant of the n xn matrix A = [aij] in the special case where n = 4 and A = '0 1 0 o 4 0 3 0 0 -1 0 -6 5\ 2 -2 o/ Do so in each of the following two ways: (a) by finding and summing the nonzero terms in the expression £ (-D*"<yi A)«iy,"■"»}„ or Y. {-l)*"UX W«i»'■■"'■•"• (where j\ jnori\ /„ is a permutation of the first n positive integers and the summation is over all such permutations); (b) by repeated expansion in terms of cofactors—use the (general) formula n n IAl = J2 aiJaiJ or IA' = 1] auau j=\ .=1 (where i or /, respectively, is any integer between 1 and n inclusive and whereof/; is the cofactor of «//) to expand |A| (in the special case) in terms of the determinants
72 13. Determinants of 3 x 3 matrices, to expand the determinants of the 3 x 3 matrices in terms of the determinants of 2 x 2 matrices, and finally to expand the determinants of the 2x2 matrices in terms of the determinants of 1 x 1 matrices. Solution, (a) |A| = (_1)^4(2,1.4.3)4(1)(_2)(_6) + (_1)^4(4.1.2,3)5(1)(3)(_6) = (-l),+0+148 + (-l)3+0+0(-90) = 48 + 90 = 138. (b) |A| = (1)(-1)2+1 = (-l)3(-6)(-l)3+2 4 0 5 3 0 -2 0 -6 0| 14 5| 3 -2 = (-l)3(-6)(-l)5[4(-l)1+1(-2) +5(-l)1+2(3)] = (-6)(-8-15) = 138. EXERCISE 6. ,Let A = {a,y} represent annxn matrix. Verify that if A is symmetric, then the matrix of cofactors (of A) is also symmetric. Solution. Let or/y represent the cofactor ofay, let A/y represent the (n — 1) x (n—1) submatrix of A obtained by striking out the /th row and the y th column (of A), and let By,- represent the {n — 1) x (n — 1) submatrix of A' obtained by striking out the yth row and the /th column of A'. Then, making use of Lemma 13.2.1 and result (2.1.1), we find that «u = (-D,+y|Al7| = (-D'+'IA^I = <-l)'+'|By,|. Moreover, if A is symmetric, then By/ = Ay/, implying that a/y = (-l)'+'|Ay/|=tfy/ and hence that the matrix of cofactors is symmetric. EXERCISE 7. Let A represent an n x n matrix. (a) Show that if A is singular, then adj(A) is singular. (b) Show that det[adj(A)l = [deKA)!""1. Solution, (a) If A is null, then it is clear that adj(A) = 0 and hence that adj(A) is singular.
13. Determinants 73 Suppose now that A is singular but nonnull, in which case A contains a nonnull row, say the /th row aj. Since A is singular, |A| = 0, and it follows from Theorem 13.5.3 that A adj (A) = 0 and hence that aj adj(A) = 0, implying (since a) is nonnull) that the rows of adj(A) are linearly dependent. We conclude that adj(A) is singular. (b) Making use of Theorems 13.3.4 and 13.5.3, Corollary 13.2.4, and result (1.9), we find that |A| |adj(A)| = |A adj(A)| = det(|A|I„) = |A|"|IW| = |A|". (S.l) If A is nonsingular, then |A| # 0, and it follows from result (S.l) that ladjtAJlHAr1. Alternatively, if A is singular, then it follows from Part (a) that adj(A) is singular and hence that |adj(A)| =0 = ^-1. EXERCISE 8. For any n x n nonsingular matrix A, A-'Ml/IADadjCA). (*) Use formula (*) to verify that, for any 2x2 nonsingular matrix A = ( n 12 ), \«21 «22/ («22 -an\ -«2i «n/' A"1 =(!/*)( "^ ""), (**) where k = a\\an — ai2«2i- Solution. Let a,-y represent the //th element of a 2 x 2 matrix A, and let or/y represent the cofactor of a,-; (/, 7 = 1,2). Then, as a special case of formula (*) [or equivalently formula (5.7)], we have that a"=«/'*)(:;; £)■ <s-2> Moreover, it follows from the very definition of a cofactor and from formulas (1.3) and (1.4) that a,, = (~l)l+la22 = «22» «21 = (~1)2+1«12 = ~«12. otn = (-l)1+2tf2i = -«2i, «22 = (-D2+2«ii =«ii, and |A| = «11^22 ~«12«21- Upon substituting these expressions in formula (S.2), we obtain formula (**) [or equivalently formula (8.1.2)].
74 13. Determinants EXERCISE 9. Let ■-KJD- (a) Compute the cofactor of each element of A. (b) Compute |A| by expanding |A| in terms of the cofactors of the elements of the second row of A, and then check your answer by expanding |A| in terms of the cofactors of the elements of the second column of A. (c) Use formula (*) of Exercise 8 to compute A-1. Solution, (a) Let a,y represent the cofactor of the //th element of A. Then, = 19. ffn =(-l)1+1 ^,3 = (-1)1+3 a22 = (-D2+2 «31 = (~1)3+1 3 1 -4 5 -1 0 2 0 1° |3 -ll 5I -1 1 = 4, = 10, = 3, <*12 = (-1)1+2 or2i=(-D2+I| «23 = (-1)2+3 1-1 1 ° 0 -4 ll 5| -1 5 = 5, = 4, «32 , = (-D3+2 2 -1 -1 1 = -1, and «33 = (-1)3+3 2 0 -1 3 = 6. (b) Expanding | A| in terms of the cofactors of the elements of the second row of A gives |A| = (-1)4 + 3(10) + 1(8) = 34. Expanding | A| in terms of the cofactors of the elements of the second column of A gives |A| = 0(5) + 3(10) + (-4)(-1) = 34. (c) Substituting from Parts (a) and (b) in formula (*) of Exercise 8 [or equivalent^ in formula (5.7)], we find that A"1 =(1/34) (19 4 3\ 5 10 -1 4 8 6/ EXERCISE 10. Let A = {a,-;} represent an n x /2 matrix (where n > 2), and let ctij represent the cofactor of A//.
13. Determinants 75 (a) Show [by for instance, making use of the result of Part (b) of Exercise 11.3] that if rank(A) = n — 1, then there exists a scalar c such that adj(A) = cxy', where x = {.vy} and y = {y,} are any nonnull /i-dimensional column vectors such that Ax = 0 and A'y = 0. Show also that c is nonzero and is expressible as c = otij/(yiXj) for any / and j such that y; ^ 0 and .\j ^ 0. (b) Show that if rank(A) < /2 - 2, then adj(A) = 0. Solution, (a) Suppose that rank(A) = n - 1. Then, det(A) = 0 and hence (according to Theorem 13.5.3) A adj(A) = adj(A)A = 0. Thus, it follows from the result of Part (b) of Exercise 11.3 that there exists a scalar c such that adj(A) = cxy' [or equivalently such that (adj A)' = cyx'] and hence such that (for all / and j) ctij=cyiXj. (S.3) Moreover, since (according to Theorem 4.4.10) A contains an (n — 1) x (n — 1) nonsingular submatrix, a,y ^ 0 for some / and j% implying that c ^ 0. And, for any / and j such that y,- ^ 0 and Xj ^ 0, we have [in light of result (S.3)] that c = ctij/lytXj). (b) If rank(A) < n — 2, then it follows from Theorem 4.4.10 that every (n — 1) x (n — 1) submatrix of A is singular, implying that a,;- = 0 for all i and j or equivalently that adj (A) = 0. EXERCISE 11. Let A represent an n x n nonsingular matrix and b an n x 1 vector. Show that the solution to the linear system Ax = b (in x) is the n x 1 vector whose yth component is |A;|/|A|, where Ay- is a matrix formed from A by substituting b for the 7th column of A (7 = 1 «). [This result is called Cramer*s rule, after Gabriel Cramer (1704- 1752).] Solution. The (unique) solution to Ax = b is expressible as A_1b. Let fy represent the / th element of b and atj the cofactor of the ijth element of A (/, j = 1 n). It follows from Corollary 13.5.4 that the jth element of A-1b is = (1/|A|) ^/wy. Clearly, the cofactor of the ijth element of A;- is the same as the cofactor of the ijth element of A (/ = 1 »). so that, according to Theorem 13.5.1, the 7'th element of A_,b is |Ay|/|A|. EXERCISE 12. Let c represent a scalar, let x and y represent n x 1 vectors, and let A represent an n x n matrix.
76 13. Determinants (a) Show that A y x7 c = c|A| - x'adKAjy. (E.1) (b) Show that, in the special case where A is nonsingular, result (E.1) can be reexpressed as A y x7 c = lAKc-tfA^y). in agreement with the more general result that, for any n xn nonsingular matrix T, n x m matrix U, m x n matrix V, and m x m matrix W, T U V W W V U Tl = |T||W-VT-1U|. (*) Solution, (a) Denote by jr,- the ith element of x, and by y,- the /th element of y. Let Ay represent the n x (n — I) submatnx of A obtained by striking out the 7th column, let A,y represent the (n — 1) x (/2 — 1) submatnx of A obtained by striking out the /th row and the 7th column, and let a,y represent the cofactor of the //th element of A. Expanding A y obtain in terms of the cofactors of the last row <i fr A y x' c = £*y(-l)n+1+ydet(Ay, y) + c(-l)2(n+1)|A|. (S.4) j Further, expanding det(Ay, y) in terms of the cofactors of the last column of (Ay, y), we obtain det(Ay, y) = £y/(-l)/+,l|Al7|. (S.5) 1 Substituting expression (S.5) in equality (S.4), we find that |x' cl = E^<-')2M+,+'+ylA0l +'IA| = ^1-^^,(-1)^^1 ij = c|A|-^yl.VyOf/y ij = c|A|-x,adj(A)y. (b) Suppose that A is nonsingular, in which case |A| £ 0. Then, using Corollary 13.5.4, result (E.1) can be reexpressed as A y x; c = |A|{c-x,[(l/|A|)adj(A)]y} = |A|(r-it'A-Iy).
13. Determinants 77 Note that this same expression can be obtained by setting T = A, U = y, V = x', and W = c in result (*) [or equivalently result (3.13)]. EXERCISE 13. Let V* represent the (n - 1) x (n - 1) submatrix of the nxn Vandermonde matrix V = /1 1 X2 \1 X„ X* 1 , (where x\, .vo x„ are arbitrary scalars) obtained by striking out the kth row and the nth (last) column (of V). Show that \Y\ = \\k\(-\)n-kY\(xk-Xi). Solution. Let V* represent the n x n matrix whose first {k — l)th rows are respectively the first (k - l)th rows of V, whose kth (n — l)th rows are respectively the (k + l)th nth rows of V, and whose /zth row is the kth row of V. Then, V* (like V) is an n x n Vandermonde matrix, and V* equals the (n — 1) x (/7- 1) submatrix of V* obtained by striking out the last row and the last column (of V*). Moreover, V can be obtained from V* by n — k successive interchanges of pairs of rows — specifically, V can be obtained from V* by successively interchanging the nth row of V* with the (n — l)th kth rows of V*. Thus, making use of Theorem 13.2.6 and of result (6.4), we find that |V| = (-i)"-*|v*| = (-l)"-A(.vA--x,)-- -*n<v* = |V*| (-1)' '(xk-xn)\\k\ EXERCISE 14. Show that, for nxn matrices A and B, adj(AB) = adj(B)adj(A). (Hint. Use the Binet-Cauchy formula to establish that the ijth element of adj(AB) equals the ijth element of adj(B)adj(A).) Solution. Let Ay represent the (n - 1) x n submatrix of A obtained by striking out the 7'th row of A, and let B,- represent the n x (n — 1) submatrix of B obtained by striking out the ith column of B. Further, let Ay, represent the (n - 1) x (n - 1) submatrix of A obtained by striking out the jth row and the sth column of A, and let B„- represent the (n - 1) x (n - 1) submatrix of B obtained by striking
78 13. Determinants out the sth row and the ith column of B. Then, application of formula (8.3) (the Binet-Cauchy formula) gives |Ay^|=^|Ay,||B„|. 5=1 implying that (-\)J+i\AjBi\ = £<-l)'+,'|Brf| (-l)J+s \AjS\. (S.6) 5=1 Note that AyB/ equals the (n — 1) x (n — 1) submatrix of AB obtained by striking out the yth row and the /th column of AB, so that the left side of equality (S.6) is the cofactor of the jith element of AB and hence is the //th element of adj(AB). Note also that (— 1),+,*|B„-| is the cofactor of the s/th element of B and hence is the /sth element of adj(B) and similarly that (-l)J+s\Ajs\ is the cofactor of the jsth element of A and hence is the sjth element of adj( A). Thus, the right side of equality (S.6) is the //th element of adj(B)adj(A). We conclude that adj(AB) = adj(B)adj(A).
14 Linear, Bilinear, and Quadratic Forms EXERCISE 1. Show that a symmetric bilinear form x'Ay (in n-dimensional vectors x and y) can be expressed in terms of the corresponding quadratic form, that is, the quadratic form whose matrix is A. Do so by verifying that x'Ay = (l/2)[(x + y)'A(x + y) - x'Ax - /Ay]. Solution. Since the bilinear form x'Ay is symmetric, we have that (l/2)[(x + y)'A(x + y) - x'Ax - /Ay] = (l/2)(x'Ax + x'Ay + y'Ax + /Ay - x'Ax - y'Ay) = (l/2)(x'Ay + /Ax) = (l/2)(x'Ay + x'Ay) = x'Ay. EXERCISE 2. Show that corresponding to any quadratic form x'Ax (in the n- dimensional vector x) there exists a unique upper triangular matrix B such that x'Ax and x'Bx are identically equal, and express the elements of B in terms of the elements of A. Solution. Let tf/y represent the //th element of A (/, j = \ n). WhenB = {tyj} is upper triangular, the conditions an = bn and at] + ayt — bjj + bji (j # i = 1 n) of Lemma 14.1.1 are equivalent to the conditions an = bn and a\j + a-ji = b{j (j > i = 1 n). Thus, it follows from the lemma that there exists a unique upper triangular matrix B such that x'Ax and x'Bx are identically equal, namely, the upper triangular matrix B = {fc,-/}, where bn = an and bij = aij+ciji (j >/ = 1 n).
80 14. Linear, Bilinear, and Quadratic Forms EXERCISE 3. Show, by example, that the sum of two positive semidefinite matrices can be positive definite. Solution. Consider the two n x n matrices I ft ft I and I ft , I. Clearly, both of these two matrices are positive semidefinite, however, their sum is the n x n identity matrix I„, which is positive definite. EXERCISE 4. Show, via an example, that there exist (nonsymmetnc) nonsingular positive semidefinite matrices. Solution. Consider the n x n upper triangular matrix /1 2 0 ... 0\ 0 1 0 ... 0 A= 0 0 1 ... 0 \0 0 0 \) For an arbitrary n-dimensional vector x = (*i, a*2, A3 xn)', we find that x'Ax = (Ai + A-2)2 + xj + • • • + xl > 0 and that x'Ax = 0 if x\ = -x2 and A3 = • • • = a„ = 0. Thus, A is positive semidefinite. Moreover, it follows from Corollary 8.5.6 that A is nonsingular. EXERCISE 5. Show, by example, that there exist an n x n positive semidefinite matrix A and an n x m matrix P (where m < n) such that P^AP is positive definite. Solution. Take A to be the n x n diagonal matrix diag(I,„, 0), which is clearly positive semidefinite, and take P to be the n x m (partitioned) matrix I ft'" J. Then, P'AP = 1/,,, which is an m x m positive definite matrix. EXERCISE 6. For an n x n matrix A and an n x m matrix P, it is the case that (1) if A is nonnegative definite, then FAP is nonnegative definite; (2) if A is nonnegative definite and rank(P) < nu then FAP is positive semidefinite; and (3) if A is positive definite and rank(P) = /?i, then FAP is positive definite. Convert these results, which are for nonnegative definite (positive definite or positive semidefinite) matrices, into equivalent results for nonpositive definite matrices. Solution. As in results (1)-(3) (of the exercise orequivalently of Theorem 14.2.9), let A represent an n x n matrix and P an n x m matrix. Upon applying results (1) - (3) with -A in place of A, we find that (T) if -A is nonnegative definite, then -FAP is nonnegative definite; (2;) if —A is nonnegative definite and rank(P) < nu then -FAP is positive semidefinite; and (3') if-A is positive definite and rank(P) = nu then -P'AP is positive definite. These three results can be restated as follows:
14. Linear, Bilinear, and Quadratic Forms 81 (T) if A is nonpositive definite, then P'AP is nonpositive definite; (2') if A is nonpositive definite and rank(P) < /w, then P'AP is negative semidefinite; and (3') if A is negative definite and rank(P) = ///, then P'AP is negative definite. EXERCISE 7. Let {Xi Xr} represent a set of matrices from a linear space V. And, let A = [ay] represent the r x r matrix whose ijth element is X,- »X; — this matrix is referred to as the Gram matrix (or the Gramian) of the set {Xj Xr) and its determinant is referred to as the Gramian (or the Gram determinant) of {X, Xr}. (a) Show that A is symmetric and nonnegative definite. (b) Show thatXi Xr are linearly independent if and only if A is nonsingular. Solution. Let Yj Y„ represent any matrices that form an orthonormal basis for V. Then, for j = 1 r, there exist scalars b\j b„j such that Xj =^,+... + ^¾. And, for/, j = 1 r. aU=XrXj Jt=l 5=1 It = ^bkibkj. *=1 Moreover, £JLi £*i&*/ is the ijth element of the rxr matrix B'B, where B is the n x r matrix whose kjth element is by (and hence where B; is the r x n matrix whose /fcth element is bu). Thus, A = B'B, and since B'B is symmetric (and in light of Corollary 14.2.14) nonnegative definite, the solution of Part (a) is complete. Now, consider Part (b). For j = 1 r, let by = (b\j bnj)'. Then, since clearly Yj Y„ are linearly independent, it follows from Lemma 3.2.4 that Xi Xr are linearly independent if and only if bi br are linearly independent. Thus, since bi br are the columns of B, X\ Xr are linearly independent if and only if rank(B) = /* or equivalently (in light of Corollary 7.4.5) if and only if rank(B'B) = r. And, since A = B'B, we conclude that X| Xr are linearly independent if and only if A is nonsingular. EXERCISE 8. Let A = {a,;} represent an n x n symmetric positive definite
82 14. Linear, Bilinear, and Quadratic Forms matrix, and let B = [by] = A ! . Show that, for i = 1 n, bn > 1 /an , with equality holding if and only if, for all j ^/, a/y = 0. Solution. Let U = (uj, U2), where uj is the ith column of I„ and U2 is the submatrix of I„ obtained by striking out the ith column, and observe that U is a permutation matrix. Define R = U'AU and S = R~!. Partition R and S as -(?' i.) - s=(:n t) [where the dimensions of both R* and S* are (72 - 1) x (n — 1)]. Then, r\\=u\Au\=au% (S.l) r' = u\ AU2 = (an, ai2 ait j_i, ai% i+\ ait n-\, ain)% (S.2) and (since S = U'BU) s1i=u/1Bui =bu. (S3) It follows from Corollary 14.2.10 that R is positive definite, implying (in light of Corollary 14.2.12) that R* is positive definite and hence (in light of Corollary 14.2.11) that R* is invertible and that R"1 is positive definite. Thus, making use of Theorem 8.5.11, we find [in light of results (S.l) and (S.3)] that ^/ = (0,7-1^1-)-1 and also that r'R-'r > 0 with equality holding if and only if r = 0. Since bu > 0 (and hence an — r'R"1 r > 0), we conclude that bn > 1 /an with equality holding if and only if r = 0 or equivalently [in light of result (S.2)] if and only if, for EXERCISE 9. Let A represent an m x n matrix and D a diagonal matrix such that A = PDQ for some matrix P of full column rank and some matrix Q of full row rank. Show that rank(A) equals the number of nonzero diagonal elements in D. Solution. Making use of Lemma 8.3.2, we find that rank(A) = rank(PDQ) = rank(DQ) = rank(D). Moreover, rank(D) equals the number of nonzero diagonal elements in D. EXERCISE 10. Let A represent an n x n symmetric idempotent matrix and V an n x n symmetric positive definite matrix. Show that rank(AVA) = tr(A).
14. Linear, Bilinear, and Quadratic Forms 83 Solution. According to Corollary 14.3.13, V = P7? for some nonsingular matrix P. Thus, making use of Corollary 7.4.5, Corollary 8.3.3, and Corollary 10.2.2, we find that rank(AVA) = rank[(PA)'PA] = rank(PA) = rank(A) = tr(A). EXERCISE 11. Show that if an n x n matrix A is such that x'Ax £ 0 for every n x 1 nonnull vector x, then A is either positive definite or negative definite. Solution. Let A represent an n x n matrix such that x'Ax ^ 0 for every n x 1 nonnull vector x. Define B = (1/2)(A + A'). Then, x'Bx = x'Ax for every n x 1 vector x. Moreover, B is symmetric, implying (in light of Corollary 14.3.5) that there exists a nonsingular matrix P and a diagonal matrix D = diag(Jj d„) such that B = P'DP. Thus, (Px)'DPx = x'Ax for every n x 1 vector x and hence (Px)'DPx ^ 0 for every n x 1 nonnull vector x. There exists no / such that d/ = 0 [since, if dt = 0, then, taking x to be the nonnull vector P_,e/, where e,- is the /th column of I„, we would have that (Px)'DPx = ejDe,- = d\ = 0). Moreover, there exists no / and j such that d\ > 0 and dj < 0 [since, if d\ > 0 and dj < 0, then, taking x to be the (nonnull) vector P_1y» where y is the n x 1 vector with /th element l/y/di and yth element l/y/-dj> we would have that (Px)'DPx = y'Dy = 1 - 1 = 0]. It follows that the n scalars d\ dn are either all positive, in which case B is (according to Corollary 14.2.15) positive definite, or all negative, in which case —B is positive definite and hence B is negative definite. We conclude (on the basis of Corollary 14.2.7) that A is either positive definite or negative definite. EXERCISE 12. (a) Let A represent an n x n symmetric matrix of rank r. Take P to be an n x n nonsingular matrix and D an n x n diagonal matrix such that A = P'DP — the existence of such matrices is guaranteed. The number, say /w, of diagonal elements of D that are positive is called the index of inertia of A (or of the quadratic form x'Ax whose matrix is A). Show that the index of inertia is well-defined in the sense that m does not vary with the choice of P or D. That is, show that, if Pi and P2 are nonsingular matrices and Di and D2 diagonal matrices such that A = P'jDjPj = P2D2P2» then D2 contains the same number of positive diagonal elements as Dj. Show also that the number of diagonal elements of D that are negative equals r — m. (b) Let A represent an n x n symmetric matrix. Show that A = P' diag(IIM, —Ir_m, 0)P for some n x n nonsingular matrix P and some nonnegative integers m and r. Show further that m equals the index of inertia of the matrix A and that r = rank(A). (c) An n x n symmetric matrix B is said to be congruent to an n x n symmetric matrix A if there exists an n x n nonsingular matrix P such that B = P'AP. (If B is congruent to A, then clearly A is congruent to B.) Show that B is congruent to
84 14. Linear, Bilinear, and Quadratic Forms A if and only if B has the same rank and the same index of inertia as A. This result is called Sylvester's law of inertia, after James Joseph Sylvester (1814-1897). (d) Let A represent an n x n symmetric matrix of rank r and with index of inertia m. Show that A is nonnegative definite if and only if m = r and is positive definite if and only if m = r = n. Solution, (a) Take Pi and P2 to be « x n nonsingular matrices and Di = [djl)\ and D2 = {d™} to be n x n diagonal matrices such that A = P'jDjPi = P2D2P2. Let mi represent the number of diagonal elements of D\ that are positive and /H2 the number of diagonal elements of D2 that are positive. Take /i, /2...., hi to be a permutation of the first n positive integers such that djl) > 0 for j = 1,2 m 1, and similarly take ki, /¾,..., k„ to be a permutation such that d™ > 0 for j = 1,2 m2- Further, take Uj to be the n x n permutation matrix whose first, second,..., /7 th columns are respectively the iith, /2th, ..., /,, th columns of |„ and U2 to be the n x n permutation matrix whose first, second nth columns are respectively the k\ th, foth Ar„th columns of I,,, and define DJ = U'jDjUi and D* = U2D2U2. Then, DJ = diag^0, dV djl)) and D| = diag^f^^ 4f>- Suppose, for purposes of establishing a contradiction, that m \ < m 2, and observe that D? =U2(P2-1),AP2-1U2 = U2(P2"1),PiD1PiP2-1U2 = U^PjVPlUiDjUjPiP^Uz = R'DJR, where R = UjPjPj'Uz. Partition the n x n matrix R as R = ( n R12 ), where R11 is of dimensions m\ x /722- Take x = [xj } to be an /^-dimensional nonnull column vector such that Ri 1 x = 0 — since (by supposition) m \ < ///2» such a vector necessarily exists. Letting y\, V2 ^ii-wi . represent the elements of the vector R21X, we find that y=i = (r2IxJD'(r21xJ Moreover, E<'v?>0, 7=1
14. Linear, Bilinear, and Quadratic Forms 85 and, since the last n — m\ diagonal elements of D* are either zero or negative, j=mi + \ These two inequalities, in combination with equality (S.4), establish the sought- after contradiction. We conclude that /«j > 7123. It can be established, via an analogous argument, that /hi < rti2. Together, those two inequalities imply that m2=m\. Consider now the number of negative diagonal elements in the diagonal matrix D. According to Lemma 14.3.1, the number of nonzero diagonal elements in D equals r. Thus, the number of negative diagonal elements in D equals r — m. (b) According to Corollary 14.3.5, there exists an/i xn nonsingular matrix P* and annxn diagonal matrix D = {<#} such that A = Pl„DP*. Take i\, /2 i„ to be any permutation of the first n positive integers such that—for some integers m and r (0 < m < r < n) — dt. > 0, for j = 1 mtdtj < 0, for j = m + 1 r, and djj = 0, for / = r + 1,..., n. Further, take U to be the nxn permutation matrix whose first, second,..., nth columns are respectively the i\th, /2 th,..., /„ th columns of I„, and define D* = U'DU. Then, D* = diag(^ll,^2 din). We find that A = P^UU'DUU'P* = (U'PJ'D^U'P*. And, taking A to be the diagonal matrix whose first m diagonal elements are y/dil* y/dil y/di^* whose (m + l)th, (m + 2)th rth diagonal elements ^ y/~d'm+i»>/~dim+i y/—dir*an^ whose last n-r diagonal elements equal one, we have that A-^A"1 =diag(Im,-Ir_m,0) and hence that A = (U'P^UU'P* = (AU'PJ'A-^A^AU'P, = P,diag(Im,-Ir_m,0)P, where P = AU'P*. Clearly, P is nonsingular. That m equals the index of inertia and that r = rank(A) are immediate consequences of the results of Part (a). (c) Suppose that B is congruent to A. Then, by definition, B = P'AP for some nxn nonsingular matrix P. Moreover, according to Corollary 14.3.5, A = Q'DQ for some nxn nonsingular matrix Q and some nxn diagonal matrix D, Thus, B = P,Q,DQP = PiDP*, where P* = QP. Clearly, P* is nonsingular. And, in light of Part (a), we conclude that B has the same rank and the same index of inertia as A.
86 14. Linear, Bilinear, and Quadratic Forms Conversely, suppose that A and B have the same rank, say r, and the same index of inertia, say m. Then, according to Part (b), A = P/ diagfl,,,, -lr-m> 0) P and B = Q' diagfl,,,, -Ir_,„, 0) Q for some n x n nonsingular matrices P and Q. Thus, (QT'BQ-1 = diagfl,,,, -Ir-„„ 0) = (P^/AP-1, and consequently B = Q'CP-^AP-'Q = P>P+, where P* = P-1Q. Clearly, P* is nonsingular. We conclude that B is congruent to A. (d) According to Part (b), A = P'diag(Il,l,-Ir_„l,0)P for some n x n nonsingular matrix P. Thus, we have as an immediate consequence of Corollary 14.2.15 that A is nonnegative definite if and only if m = r and is positive definite if and only if m =/• = «. EXERCISE 13. Let A represent an n x n symmetric nonnegative definite matrix of rank r. Then, there exists an n x r matrix B (of rank r) such that A = BB'. Let X represent any n x m matrix (where //* > r) such that A = XX'. (a) Show that X = PBX. (b) Show that X = (B, 0)Q for some orthogonal matrix Q. Solution, (a) It follows from Corollary 7.4.5 that C(X) = C( A) = C(B), implying (in light of Corollary 12.3.6) that Px = Pb- Thus, making use of Part (1) of Theorem 12.3.4, we find that X = PXX = PBX. (b) Since rank(B'B) = rank(B) = /\ B'B (which is of dimensions /• x r) is invertible. Thus, it follows from Part (a) that X = PBX = BfB'Br'B'X = BQ,, (S.5) where Qj = (B'B^B'X. Moreover, 0,(¾ = (B'Br'B'XX'BfB'Br1 = (B'B^B'ABfB'Br1 = (B'BrVBB'BBfB'Br1 =1, so that the rows of the r x m matrix Qj are orthonormal (with respect to the usual inner product). It follows from Theorem 6.4.5 that there exists an (/h - r) x m matrix Q2 whose rows, together with the rows of Qi, form an orthonormal (with respect to the usual
14. Linear, Bilinear, and Quadratic Forms 87 inner product) basis for Wn. Take Q = I q1 ) • Then, clearly, Q is orthogonal. Further, (B, 0)Q = BQ,, implying, in light of result (S.5), that X = (B, 0)Q. EXERCISE 14. Show that if a symmetric matrix A has a nonnegative definite generalized inverse, then A is nonnegative definite. Solution. Suppose that the symmetric matrix A has a nonnegative definite generalized inverse, say G. Then, A = AGA = A'GA, implying (in light of Theorem 14.2.9) that A is nonnegative definite. EXERCISE 15. Suppose that annxn matrix A has an LDU decomposition, say A = LDU, and let d\, d2t • •»dn represent the diagonal elements of the diagonal matrix D. Show that \A\ = dld2---dn. Solution. Making use of Theorem 13.2.11 and of Corollary 13.1.2, we find that |A| = ILDUI = |LD| = |D| = dxd2 -dn. EXERCISE 16. (a) Suppose that an n x n matrix A (where n > 2) has a unique LDU decomposition, say A = LDU, and let d\t d2 d„ represent the first, second nth diagonal elements of D. Show that 4/0(/ = 1,2 n — 1) and that d„ / 0 if and only if A is nonsingular. (b) Suppose that annxn (symmetric) matrix A (where n > 2) has a unique U'DU decomposition, say A = U'DU, and let 4» ^2* • • -»4 represent the first, second nth diagonal elements of D. Show that d\ / 0 (i = 1,2 n — 1) and that dn / 0 if and only if A is nonsingular. Solution. Let us restrict attention to Part (a)—Part (b) can be proved in essentially the same way as Part (a). Suppose — for purposes of establishing a contradiction — that, for some / (1 < i < n — 1), d\ = 0. Take L* to be a unit lower triangular matrix and U* a unit upper triangular matrix that are identical to L and U, respectively, except that, for some ;' (j > i) the //th element of U* differs from the //th element of U and/or the jith element of L* differs from the jilh element of L. Then, according to Theorem 14.5.5, A = L*D*U* is an LDU decomposition of A. Since this decomposition differs from the supposedly unique LDU decomposition A = LDU, we have arrived at the desired contradiction. We conclude that 4/0(/ = 1 n - 1). And, since (in light of Lemma 14.3.1) A is nonsingular if and only if all n diagonal elements of D are nonzero, we further conclude that A is nonsingular if and only if4/0. EXERCISE 17. Suppose that an n x n (symmetric) matrix A has a unique U'DU decomposition, say A = U'DU. Use the result of Part (b) of Exercise 16 to show that A has no LDU decompositions other than A = U'DU.
88 14. Linear, Bilinear, and Quadratic Forms Solution. Let us restrict attention to the case where n > 2 — if n = 1, then it is clear that A has no LDU decompositions other than A = U'DU. The result of Part (b) of Exercise 16 implies that the first n -1 diagonal elements of D are nonzero. We conclude, on the basis of Theorem 14.5.5, that A has no LDU decompositions other than A = U'DU. EXERCISE 18. Show that if a nonsingular matrix has an LDU decomposition, then that decomposition is unique. Solution. Any 1 x 1 matrix (nonsingular or not) has a unique LDU decomposition, as discussed in Section 14.5b. Consider now a nonsingular matrix A of order i! > 2 that has an LDU decomposition, say A = LDU. Let An, Ln, Un, and Di represent the (n - l)th-order leading principal submatrices of A, L, U, and D, respectively. Then, according to Theorem 14.5.3, An = LnDjUn is an LDU decomposition of An- Since A is nonsingular, D is nonsingular, implying that Dj is nonsingular and hence that An is nonsingular. We conclude, on the basis of Corollary 14.5.6, that A = LDU is the unique LDU decomposition of A. EXERCISE 19. Let A represent an n x n matrix (where n > 2). By for instance using the results of Exercises 16, 17, and 18, show that if A has a unique LDU decomposition or (in the special case where A is symmetric) a unique U'DU decomposition, then the leading principal submatrices (of A) of orders 1,2 »—1 are nonsingular and have unique LDU decompositions. Solution. In light of the result of Exercise 17, it suffices to restrict attention to the case where A has a unique LDU decomposition, say A = LDU. For i = 1,2,..., n — 1, let A/, L,-, U,\ and D/ represent the /th-order leading principal submatrices of A, L, U, andD, respectively. Then, according to Theorem 14.5.3, an LDU decomposition of A/ is A,- = L,D,U/, and, according to the result of Exercise 16, D/ is nonsingular. Thus, A/ is nonsingular and, in light of the result of Exercise 18, has a unique LDU decomposition. EXERCISE 20. (a) Let A = [aij] represent an m x n nonnull matrix of rank /*. Show that there exist an m x m permutation matrix P and an n x n permutation matrix Q such that where Bj i is an r x r nonsingular matrix whose leading principal submatrices (of orders 1,2 r — 1) are nonsingular. (b) Let B = f R _ *" J represent any m x n nonnull matrix of rank /■ such that Bn is an rxr nonsingular matrix whose leading principal submatrices (of orders 1,2 /■ — 1) are nonsingular. Show that there exists a unique decomposition of
14. Linear, Bilinear, and Quadratic Forms 89 B of the form b=(lOD(U,,U2)- where Lj is an r x r unit lower triangular matrix, Uj is an /• x r unit upper triangular matrix, and D is an r x r diagonal matrix. Show further that this decomposition is suchthatBn = LiDUi is the unique LDU decomposition of Bn,D is nonsingular, L2 = BziI^D"1, and U2 = D^L^'B^ . Solution, (a) The matrix A contains r linearly independent rows, say rows /|, h ir- For k = 1 r, denote by A* the k x n matrix whose rows are respectively rows i\. iS,..., i* of A. There exists a subset 71,72 jr of the first n positive integers such that, for k = 1 r, the matrix, say A£, whose columns are respectively columns 71,/2 jk of Ajt, is nonsingular. As evidence of this, let us outline a recursive scheme for constructing such a subset. Row ii of A is nonnull, so that 71 can be chosen in such a way that 0,-,7, ^ 0 and hence in such a way that Aj = (fl;,y,) is nonsingular. Suppose now that 71, 72 7a—1 have been chosen in such a way that A*, Aj A£_, are non- singular. Since A£_, is nonsingular, columns 71,72 7Jt-i of A* are linearly independent, and, since rank (A*) = kt A* has a column that is not expressible as a linear combination of columns 7*1,72....»A-1- Thus, it follows from Corollary 3.2.3 that /jt can be chosen in such a way that A£ is nonsingular. Take P to be any m x m permutation matrix whose first r rows are respectively rows i], /2,..., ir of I,„, and take Q to be any n x n permutation matrix whose first r columns are respectively columns 71,72 7V of ln- Then, —ft $• where Bn = A* is a nonsingular matrix whose leading principal submatrices (of orders 1,2,..., r -1) are respectively the nonsingular matrices A J, A| A*_,. (b) Clearly, showing that B has a unique decomposition of the form specified in the exercise is equivalent to showing that there exist a unique unit lower triangular matrix Li, a unique unit upper triangular matrix Ui, a unique diagonal matrix D, and unique matrices L2 and U2 such that Bli=LiDU,, B12 = L!DU2, B21 =L2DUi, and B22 = L2DU2. It follows from Corollary 14.5.7 that there exists a unique unit lower triangular matrix Li, a unique unit upper triangular matrix Ui, and a unique diagonal matrix D such that Bn = L1DU1 — by definition, Bn = LjDUj is the unique LDU decomposition of Bn- Moreover, D is nonsingular (since Bn = LjDUj is non- singular). Thus, there exist unique matrices L2 and U2 such that B21 = L2DU1
90 14. Linear, Bilinear, and Quadratic Forms and B12 = LjDU2 , namely, L2 = feU^D-1 and U2 = D~lLY lB\2 . Finally, it follows from Lemma 9.2.2 that B22 = B21B7/BJ2 = L2DUi(L1DUi)-1LiDU2 = L2DUiU71D1L71LiDU2 = L2DU2 . EXERCISE 21. Show, by example, that there exist nxn (nonsymmetric) positive semidefinite matrices that do not have LDU decompositions. Solution. Let Consider the quadratic form x'Ax in x. Partitioning x as x = I l J, where X2 is of dimensions (n — 1) x 1, we find that X'AX = - Yi (l'x2) + .VI (Xjl) + X2IX2 = X2X2. Thus, x'Ax > 0 for all x, with equality holding when, for example, x\ = 1 and X2 = 0, so that x'Ax is a positive semidefinite quadratic form and hence A is a positive semidefinite matrix. Moreover, since the leading principal submatrix of A of order two is I . . J and since —\£ C(0), it follows from Part (2) of Theorem 14.5.4 that A does not have an LDU decomposition. EXERCISE 22. Let A represent an n x n nonnegative definite (possibly non- symmetric) matrix that has an LDU decomposition, say A = LDU. Show that the diagonal elements of the diagonal matrix D are nonnegative. Solution. Consider the matrix B = DU(L-1)'. Since (in light of Corollary 8.5.9) (L-1)' — like U — is unit upper triangular, it follows from Lemma 1.3.1 that the diagonal elements of B are the same as the diagonal elements, say d\ d„. of D. Moreover, B = L-'tLDUXL-1)' = L-'Aar1)', implying (in light of Theorem 14.2.9) that B is nonnegative definite. We conclude — on the basis of Corollary 14.2.13 — that d\ d„ are nonnegative. EXERCISE 23. Let A represent an m x k matrix of full column rank. And, let A = QR represent the QR decomposition of A; that is, let Q represent the unique m x k matrix whose columns are orthonormal with respect to the usual inner product and let R represent the unique k x k upper triangular matrix with positive diagonal elements such that A = QR. Show that A'A = R'R (so that A'A = R'R is the Cholesky decomposition of A'A).
14. Linear. Bilinear, and Quadratic Forms 91 Solution. Since the inner product with respect to which the columns of Q are orthonormal is the usual inner product, Q'Q = I*, and consequently A'A = R'Q'QR = R'R. EXERCISE 24. Let A represent an m x k matrix of rank r (where r is possibly less than k). Consider the decomposition A = QRi, where Q is an m x /• matrix with orthonormal columns and R| is an r x k submatrix whose rows are the r nonnull rows of a k x k upper triangular matrix R having r positive diagonal elements and n — r null rows. (Such a decomposition can be obtained by using the results of Exercise 6.4 — refer to Exercise 6.5.) Generalize the result of Exercise 23 by showing that if the inner product with respect to which the columns of Q are orthonormal is the usual inner product, then A'A = R'R (so that A'A = R'R is the Cholesky decomposition of A'A). Solution. Suppose that the inner product with respect to which the columns of Q are orthonormal is the usual inner product. Then, Q'Q = Ir. Thus, recalling result (2.2.9), we find that A'A = R',Q'QRi = R',Ri = R'R. EXERCISE 25. Let A = {fl,y} represent an n x n matrix that has an LDU decomposition, say A = LDU. And, define G = tJ-'D""!/"1 (which is a generalized inverse of A). (a) Show that G = D-L"1 + (I - U)G = IT'D" + G(I-L). (b) For / = 1 /i, let <// represent the /th diagonal element of the diagonal matrix D; and, for /, j = 1 n, let £,;, «,y, and g,j represent the ijth elements of L, U, and G, respectively. Take D~ = diagfc/* d*), where d* = 1/</,-, if d\ £ 0, and d* is an arbitrary scalar, if d\ = 0. Show that n n k=i+l k=i+l and that & j = (where the degenerate sums T!k=n+i Sikhi and T!k=n+i u'k8ki are to be interpreted as 0). £ gikekj, k=j+l n J2 U'tk8kJ> k=i+\ for j < i , for j > i (E.2a) (E.2b)
92 14. Linear, Bilinear, and Quadratic Forms (c) Devise a recursive procedure that uses the formulas from Part (b) to generate a generalized inverse of A. Solution, (a) Clearly, D~L~l + (1- U)G = D-IT1 + G - UG = D"^1 + G - UU^D-IT1 = G. Similarly, U-1D-+G(I-L) = U~1D-+G-GL = ^^-+0-^^-^^ = ^ (b) Since L and U are unit triangular, their diagonal elements equal 1 (and the diagonal elements of I - L and I - U equal 0). Thus, it follows from Part (a) that gij = and similarly that n > Jfe-i+1 if ./ = 1 if j > i 8ij = d*- J^ gikhi, if 7=*\ k=i+l Jt=/+i if j < i. (c) The formulas from Part (b) can be used to generate a generalized inverse of A in n steps. During the first step, the /2 th diagonal element gnn is generated from the formulagnn = tf*, and then£„_!,„ gi„ and£„,„_! g„j (theoff-diagonal elements of the nth column and row of G) are generated recursively using formulas (E.2b) and (E.2a), respectively. During the (n - s + l)th step (n - 1 < s < 2), the sth diagonal element gS5 is generated from gs+\ts 8ns or alternatively from gSts+i gsn using result (E.1), and then^_i.5 g\s and ^,^_i g5\ are generated recursively using formulas (E.2b) and (E.2a), respectively. During the /2th (and final) step, the first diagonal element gi i is generated from the last n — 1 elements of the first column or row using result (E.1). EXERCISE 26. Verify that a principal submatrix of a skew-symmetric matrix is skew-symmetric. Solution. Let B = [bij) represent the r x r principal submatrix of an n x n skew- symmetric matrix A = {atj) obtained by striking out all of its rows and columns except the fc|th, *2th krth rows and columns (where k\ < k2 < ..., kr). Then, for/, j = 1 /\ bji =akjk, = ~Ok,kj = -bij.
14. Linear. Bilinear, and Quadratic Forms 93 Since bj\ is the /;th element of B' and -fc,-; the //th element of -B, we conclude thatB' = -B. EXERCISE 27. (a) Show that the sum of skew-symmetric matrices is skew- symmetric. (b) Show that the sum Ai + Ao + • • • + A* of n x n nonnegative definite matrices Ai, A2 A* is skew-symmetric if and only if Ai, A2 Ajt are skew-symmetric. (c) Show that the sum Ai + A2 H 1- Ajt of n x n symmetric nonnegative definite matrices Ai, A2 A* is a null matrix if and only if Ai, A2 A* are null matrices. Solution, (a) Let Ai, A2 A* represent n x n skew-symmetric matrices. Then, so that 5^,- A,- is skew-symmetric. (b) If the nonnegative definite matrices A\, A2 Ajt are skew-symmetric, then it follows from Part (a) that their sum £,• A/ is skew-symmetric. Conversely, suppose that £. A,- is skew-symmetric. Let d,j represent the yth diagonal element of A/ (i = 1,..., k\ j = 1 /1). Since (according to the definition of skew-symmetry) the diagonal elements of £,- A/ equal zero, we have that dlj+d2j + -' + dkj=0 (j = 1,..., n). Moreover, it follows from Corollary 14.2.13 that d\j > 0, d2j>0, ..., dkj > 0, leading to the conclusion that d\j, d2j dkj equal zero 0 = 1 n). Thus, it follows from Lemma 14.6.4 that Aj, A2 A* are skew-symmetric. (c) Since (according to Lemma 14.6.1) the only n x n symmetric matrix that is skew-symmetric is the n x n null matrix, Part (c) is a special case of Part (b). EXERCISE 28. (a) Let Ai, A2 A* represent n x n nonnegative definite matrices. Show that tr(£/=i A,) > 0, with equality holding if and only if J^_, A/ is skew-symmetric or equivalently if and only if Ai, A2 Ajt are skew-symmetric. [Note. That 5Z/=i A/ being skew-symmetric is equivalent to Aj, A2 Ajt being skew-symmetric is the result of Part (b) of Exercise 27.] (b)Let Ai, A2 Ajt represent n xn symmetric nonnegative definite matrices. Show that tr(£/=i A/) > 0, with equality holding if and only if £*=1 A,- = 0 or equivalently if and only if Ai, A2 A* are null matrices. Solution, (a) According to Corollary 14.2.5, J^=\ A/ is nonnegative definite. Thus, it follows from Theorem 14.7.2 that tr(£f=1 A,) > 0, with equality holding
94 14. Linear, Bilinear, and Quadratic Forms if and only if 5Zf=i A/ is skew-symmetric or equivalently [in light of the result of Part (b) of Exercise 27] if and only if Aj, A2,..., A* are skew-symmetric. (b) Part (b) follows from Part (a) upon observing (on the basis of Lemma 14.6.1) that a symmetric matric is skew-symmetric if and only if it is null. EXERCISE 29. Show, via an example, that (for 77 > 1) there exist n x n (non- symmetric) positive definite matrices A and B such that tr(AB) < 0. Solution. Take A = [ay} to be an n x n matrix such that = 1, for j = 1, UU = 2, for y = J + l, for ; = 1 -1 (/ = 1 n), and take B = A. That is, take B = A = / 1 -2 V-2 2 2 1 2 -2 1 Then, (1/2)(A+A') = I„, which is a positive definite matrix, implying (in light of Corollary 14.2.7) that A is positive definite. Moreover, all n diagonal elements of AB equal 1 — 4(/1 — 1). which (for n > 1) is a negative number. Thus, tr(AB) < 0. EXERCISE 30. (a) Show, via an example, that (for n > 1) there exist n x n symmetric positive definite matrices A and B such that the product AB has one or more negative diagonal elements (and hence such that AB is not nonnegative definite). (b) Show, however, that the product of two n x n symmetric positive definite matrices cannot be nonpositive definite. Solution, (a) Take A = diag(An, I«-2> and B = diag(Bn, I„_2), where AH-(_; "9 ana „,_(« ;). and consider the quadratic forms x'Ax and x'Bx in the /i-dimensional vector x = Ui, X2 xH)'. We find that x'Ax= 2[xi-(l/2).v2]2+(3/2).v; + .v5+.vJ + ---+.v,7, x'Bx = 12[.v, + (1/4) .v2]2 + (1/4) xl + xj +.vj + • - - +.v;. Clearly, x'Ax > 0 with equality holding only if x = 0, and similarly x'Bx > 0 with equality holding only if x = 0. Thus, the quadratic forms x'Ax and x'Bx are
14. Linear. Bilinear, and Quadratic Forms 95 positive definite, and hence, by definition, the matrices A and B of the quadratic forms are positive definite. Consider now the product AB. We find that AB = diag(A 11B11, I„_2) and that AiiBn = ( J, thereby revealing that the second diagonal element of AB equals the negative number — 1. (b) Let A and B represent n x n symmetric positive definite matrices. Suppose, for purposes of establishing a contradiction, that AB is nonpositive definite. Then, by definition, -AB is nonnegative definite, implying (in light of Theorem 14.7.2) that tr(-AB) > 0 and hence that tr(AB) = -tr(-AB) < 0. However, according to Theorem 14.7.4, tr(AB) > 0. Thus, we have arrived at the sought-after contradiction. We conclude that AB cannot be nonpositive definite. EXERCISE 31. Let A = [atj) and B = {£,;} represent n x n matrices, and take C to be the n x n matrix whose ijth element c,-y = atjbij is the product of the //th elements of A and B — C is the so-called Hadamard product of A and B. Show that if A is nonnegative definite and B is symmetric nonnegative definite, then C is nonnegative definite. Show further that if A is positive definite and B is symmetric positive definite, then C is positive definite. [Hint. Taking x = (am,..., x,,)' to be an arbitrary n x 1 vector and F = (f| f„) to be a matrix such that B = FF, begin by showing that x'Cx = tr( AH), where H = G'G withG = Uifj .v„f/,).] Solution. Suppose that B is symmetric nonnegative definite. Then, according to Corollary 14.3.8, there exists a matrix F = (fi f„) such that B = FT. Let x = (a-i .v„)' represent an arbitrary n-dimensional column vector, let G = (.vjfi xnf„), and let H = G'G. Then, the //th element of H is htj = (XifiYxjfj = XiXjf[fj = XjXjbjj. Thus, x'Cx = ^CijXiXj = ^aijbijxixj = ^aijlifj i, j «\ j ». J = ^a,v/iy7=tr(AH). ». j Clearly, the matrix H is symmetric nonnegative definite, implying (in light of Theorem 14.7.6) that if A is nonnegative definite, then tr(AH) > 0 and consequently x'Cx > 0. Consider now the special case where B is symmetric positive definite. In this special case, rank(F) = rank(B) = n, implying that the columns of F are linearly independent and hence nonnull. Thus, unless x = 0, G is nonnull and hence H is
96 14. Linear, Bilinear, and Quadratic Forms nonnull. It follows (in light of Theorem 14.7.4) that if A is positive definite, then, unless x = 0, tr(AH) > 0 and consequently x'Cx > 0. We conclude that if A is nonnegative definite and B is symmetric nonnegative definite, then C is nonnegative definite and that if A is positive definite and B is symmetric positive definite, then C is positive definite. EXERCISE 32. Let Aj, A2 A* and Bi, B2 B* represent n x n symmetric nonnegative definite matrices. Show that tr(£JLj A,B,) > 0, with equality holding if and only if, for / = 1,2 k, A,B,- = 0, thereby generalizing and the results of Part (b) of Exercise 28. Solution. According to Corollary 14.7.7, tr(A,B,) > 0 (/ = 1,2,...,k). Thus, *(£A,B,)=£tr(A,B,)>0, with equality holding if and only if, for 1 = 1,2 k, tr(A,B,) = 0 or equivalent^ (in light of Corollary 14.7.7) if and only if, for i = 1,2 k, A,B, = 0. EXERCISE 33. Let A represent a symmetric nonnegative definite matrix that has been partitioned as where T (and hence W) is square. Show that VT~U and UW~V are symmetric and nonnegative definite. Solution. According to Lemma 14.8.1, there exist matrices R and S such that T = R'R, U = R'S, V = S'R, W = S'S. Thus, making use of Parts (6) and (3) of Theorem 12.3.4, we find that VT~U = S'R(R'R)-R'S = S'PrS = S'PrPrS = S'PrPrS = (PrS/PrS and similarly that UWV = R'SfS'SrS'R = R'PSR = R'PsPsR = R'PgPsR = (PsR/PsR We conclude that VT~U and UW~V are symmetric and (in light of Corollary 14.2.14) nonnegative definite. EXERCISE 34. Show, via an example, that there exists an (m + n) x (m + n) (T U\ v w),w^ere T is of dimensions m xnuY/ of dimensions n x n, U of dimensions m x /?, and V of
14. Linear, Bilinear, and Quadratic Forms 97 dimensions nxm, for which C(U) £ C(T) and/or ft(V) £ ft(T), the expression rank(T) +rank(W - VT~U) does not necessarily equal rank(A), and the formula /T-+T-UQ-VT- -T~UQ-\ V -Q-VT- Q- )' (*} where Q = W — VT~U, does not necessarily give a generalized inverse of A. Solution. Consider the matrix A = ( J, where T = 0, U = JlfI„, V = -JU, and W = I„. Clearly, C(U) £ C(T) and ft(V) £ ft(T). Moreover, (1/2) (A + A') = I J. _ I (which is a positive semidefinite matrix), so that (according to Corollary 14.2.7) A is positive semidefinite. Now, take T~ = 0, in which case the Schur complement of T relative to T~ is W - VT~U = I„. Using Theorem 9.6.1, we find that rank(A) = n + rank(T - UW1 V) = « + rank(JlfII,Jwm) = « + rank(«J,niM) = « + 1. However, rank(T) + rank(W - VT~U) = rank(0) + rank(I„) = n. Thus, rank(A) # rank(T) + rank(W - VT~U). Further, the matrix obtained by applying formula (*) [or equivalently formula (9.6.2)] is (I J J. Since the matrix ( ft f J is not a generalized inverse of A. EXERCISE 35. Show, via an example, that there exists an (m -f /i) x (m + (T U\ it' w I. where T is of dimensions m x m, V of dimensions m x n, and W of dimensions n x /i, such that T is nonnegative definite and (depending on the choice of T~) the Schur complement W - U'T-U of T relative to T~ is nonnegative definite, but A is not nonnegative definite. Solution. Consider the symmetric matrix A = I -., w I, where T = 0, U = Jmn, and W = 1«. And, take T~ = 0, in which case the Schur complement of T relative toT-isW-U,T-U = Iw.
98 14. Linear, Bilinear, and Quadratic Forms Then, clearly, T is nonnegative definite, and the Schur complement of T relative to T~ is nonnegative definite. However, A is not nonnegative definite, as is evident from Corollary 14.8.2. EXERCISE 36. An n x n matrix A = [ajj) is said to be diagonally dominant if, for i = 1,2,..., n, \an\ > £'j=i (^,) jtffy |. (In the degenerate special case where n = 1, A is said to be diagonally dominant if it is nonnull). (a) Show that a principal submatrix of a diagonally dominant matrix is diagonally dominant. (b) Let A = [oij] represent an n x n diagonally dominant matrix, partition A as A = ( , V ) [so that An is of dimensions {n — 1) x {n — 1)], and let V D ann) C = An - (l/fl„„)ab' represent the Schur complement of ann. Show that C is diagonally dominant. (c) Show that a diagonally dominant matrix is nonsingular. (d) Show that a diagonally dominant matrix has a unique LDU decomposition. (e) Let A = {ajj} represent an n x n symmetric matrix. Show that if A is diagonally dominant and if the diagonal elements a\ \, #22 am of A are all positive, then A is positive definite. Solution, (a) Let A = [ay] represent an n x n diagonally dominant matrix, and let B = [bit) represent the /« x m principal submatrix obtained by striking out all of the rows and columns of A except the jjth, /2th /wth rows and columns (where i\ < /"2 < • ■ • < /in). Then, for k = 1,2, —, nu n m m \bkk\ = \aikik\> J2 ^- H Ku\= J2 i*«i- j=\U&k) e=\(t-£k) i=\tf&) Thus, B is diagonally dominant. (b) For /, j = 1,2 n — 1, let qy represent the //th element of C. By definition, cu = °u ~ aina,ij/ann. Then,fori = 1,2 n- I, 11-1 11-1 5^ \Cij\ < ^2 (\0U\ + \a»'anj/a't't\) /1 11-1 11-1 < Vhi I ~ |«in I + Yl \a»>a»jA7"" 1
14. Linear, Bilinear, and Quadratic Forms 99 ii-1 < k.il - kin I + ^2 \ai„a„j/ann\ j=\ (since \an\ = \an -ainani/onn +ainani/ann\ < {an - ainani/ann\ + \ainani/afm\ = \cn\ + \aina„i/ann\) ii-I = Iq/I ~ kill + kin 15^ \a„j/ann\ y=i < \cii\-\ain\ + \aht\ (since Z"Z\ \onj/a,uA = Z"Z\ \anj\/\a„„\ < \am\/\ann\ = 1) = |Q/|. (c) The proof is by mathematical induction. Clearly, any I x 1 diagonally dominant matrix is nonsingular. Suppose now that any (n — 1) x (n — 1) diagonally dominant matrix is nonsingular, and let A = {ay} represent an arbitrary n x n diagonally dominant matrix. It suffices to show that A is nonsingular. Partition A as A = ( J,1 ) [so that An is of dimensions (n — 1) x (n -1)]. \ O Ann/ Since A is diagonally dominant, ann £ 0. Let C = An - (l/fl„„)ab'. It follows from Part (b) that C is diagonally dominant, so that by supposition C [which is of dimensions (n — I) x (n — 1)] is nonsingular. Based on Theorem 8.5.11, we conclude that A is nonsingular. (d) It follows from Part (a) that every principal submatrix of a diagonally dominant matrix is diagonally dominant and hence—in light of Part (c)—nonsingular. We conclude—on the basis of Corollary 14.5.7—that a diagonally dominant matrix has a unique LDU decomposition. (e) The proof is by mathematical induction. Clearly, any 1 x 1 diagonally dominant matrix with a positive (diagonal) element is positive definite. Suppose now that any (w — 1) x (n — 1) symmetric diagonally dominant matrix with positive diagonal elements is positive definite. Let A = {a,y} represent an n x n symmetric diagonally dominant matrix with positive diagonal elements. It suffices to show that A is positive definite. Partition A as A = ( \l a ) [so that An is of dimensions («-l)x («-!)], \ 3 Aim/ and let C = An - (l/^niI)aa; represent the Schur complement of ann. It follows from Part (b) that C is diagonally dominant. Moreover, the /th diagonal element ofCis an —ai„a„j/a,w > an — |«,„| \a„i/a„n\ >ou-\a!n\ >0
100 14. Linear, Bilinear, and Quadratic Forms (i = 1,2 n — 1). Thus, by supposition, C [which is symmetric and of dimensions (/2 — 1) x (/i — 1)] is positive definite. Based on Corollary 14.8.6, we conclude that A is positive definite. EXERCISE 37. Let A = [ay] represent annx« symmetric positive definite matrix. Show that det(A) < fj"=i on, with equality holding if and only if A is diagonal. Solution. That det(A) = Yl'i=i °u if A is diagonal is an immediate consequence of Corollary 13.1.2. Thus, it suffices to show that if A is not diagonal, then det(A) < FI/=i au- This is accomplished by mathematical induction. Consider a 2 x 2 symmetric matrix A/aii *n\ V*12 022/ that is not diagonal. (Every 1 x 1 matrix is diagonal.) Even if A is not positive definite, we have that det(A) = ana22 - a\2 < anazi. Suppose now that, for every (« — 1) x (« — 1) symmetric positive definite matrix that is not diagonal, the determinant of the matrix is less than the product of its diagonal elements, and consider the determinant of an n x « symmetric positive definite matrix A = [mj} that is not diagonal (where n > 3). Partition A as A=(Ar a ) \a Qnn) [where A* is of dimensions (n — 1) x (n — 1)]. Then, in light of the discussion of Section 14.8a, it follows from Theorem 13.3.8 that |A| = |A*| (aw„ -a'A-'a). (S.6) And, it follows from Corollary 14.8.6 and Lemma 14.9.1 that |A*| > 0 and ann - a'A"1 a > 0. In the case where A* is diagonal, we have (since A is not diagonal) that a # 0, implying (since A"1 is positive definite) that a'A"1 a > 0 and hence that an„ > ann — a'A^1 a, so that [in light of result (S.6)] /j-i ii |A| < am\\*\ = am Y\ an = J~| ait. /=1 i=l In the alternative case where A* is not diagonal, we have that a'A~l a > 0, implying that ann > a„„ — a'A^1 a, and we have, by supposition, that |A*| < Yi'lZi au*s0 that n-l /j-1 n |A| < (ann - a'A^1 a) J~| an < a„„ J~| an = J~|an. /=i i=i i=i
14. Linear, Bilinear, and Quadratic Forms 101 Thus, in either case, |A| < fEU ««• EXERCISE 38. Let A = (a bA, where a, b, c, and d are scalars. i-J) (a) Show that A is positive definite if and only if a > 0, d > 0, and | b + c \ /2<y/ad. (b) Show that, in the special case where A is symmetric (i.e., where c = b), A is positive definite if and only if a > 0, d > 0, and | b \ < yfad. Solution, (a) Let B = (1/2)(A + A') = ((fc;c)/2 »+//*). Observe that det(B) = ad - [(b + c)/2]2 (S.7) and (in light of Corollary 14.2.7) that A is positive definite if and only if B is positive definite. Suppose that A is positive definite (and hence that B is positive definite). Then, it follows from Corollary 14.2.13 that a > 0 and d > 0, and [in light of equality (S.7)] it follows from Lemma 14.9.1 that ad - [(b + c)/2]2 > 0, or equivalent^ that [(b + c)/2]2 < ad, and hence that \b + c\/2 < yfad. Conversely, suppose that a > 0, d > 0, and \b + c\/2 < y/ad, in which case [{b + c)/2]2 < ad or equivalently [in light of equality (S.7)] that det(B) > 0. Then, it follows from Theorem 14.9.5 that B is positive definite and hence that A is positive definite. (b) Part (a) follows from Part (b) upon observing that, in the special case where c = b, the condition \b + c\/2 < y/ad simplifies to the condition |&| < *Jad. EXERCISE 39. By, for example, making use of the result of Exercise 38, show that if an n x n matrix A = {ay} is symmetric positive definite, then, for j £i = 1 n, \aij\ < y/aaajj < maxfo/, ajj). Solution. Suppose that A is symmetric positive definite. Clearly, the 2 x 2 matrix (a,i ai] ) is a principal submatrix of A and hence (in light of Corollary 14.2.12) aji ajjj is symmetric positive definite. Thus, it follows from Part (b) of Exercise 38 that an > 0, ajj > 0, and |fliy| < Janajj. Moreover, if an > ajj, then y/oiioj] < yja\ = an = max(a/,-, fl;y); and similarly if a,-,- < ajj, then Jauajj < y/ajj = ajj = max(flfI-,ajj).
102 14. Linear, Bilinear, and Quadratic Forms EXERCISE 40. Show, by example, that it is possible for the determinants of both leading principal submatrices of a 2 x 2 symmetric matrix to be nonnegative without the matrix being nonnegative definite and that, for n > 3, it is possible for the determinants of all n leading principal submatrices of an n x /2 symmetric matrix to be nonnegative and for the matrix to be nonsingular without the matrix being nonnegative definite. Solution. Consider the 2 x 2 symmetric matrix [ n _. J. The determinants of both of its leading principal submatrices are zero (and hence nonnegative), but it is obviously not nonnegative definite. Next, consider the 3 x 3 symmetric matrix /0 0 1\ A,= 0 -1 0 . V o i/ By, for example, expanding |A*| in terms of the cofactors of the three elements of the first row of A*, we find that |A*| = 1. Thus, the determinants of the leading principal submatrices of A* (of orders 1,2, and 3) are 0,0, and 1, respectively, all of which are nonnegative; and A* is nonsingular. However, A* is not nonnegative definite (since, e.g., one of its diagonal elements is negative). Finally, for n > 4, consider the n x n symmetric matrix Clearly, the leading principal submatrices of A of orders 1,2, and 3 are the same as those of A*, so that their determinants are 0,0, and 1, respectively. Moreover, it follows from results (13.3.5) and (13.1.9) that the determinants of all of the leading principal submatrices of A of order 4 or more equal |A*| and hence equal 1. Thus, the determinants of all n leading principal submatrices of A are nonnegative, and A is nonsingular. However, A is not nonnegative definite (since, e.g., one of its diagonal elements is negative). EXERCISE 41. Let V represent a subspace of 11" x l of dimension r (where r > 1) .TakeB = (bi, b2 br)tobe any n xr matrix whose columns bj, b2 br form a basis for V, and let L represent any left inverse of B. Let g represent a function that assigns the value x * y to an arbitrary pair of vectors x and y in V. (a) Let / represent an arbitrary inner product for llr* \ and denote by s • t the value assigned by / to an arbitrary pair of r-dimensional vectors s and t. Show that g is an inner product (for V) if and only if there exists an / such that (for all x and y in V) x*y = (Lx)*(Ly). (b) Show that g is an inner product (for V) if and only if there exists an /• x r
14. Linear, Bilinear, and Quadratic Forms 103 symmetric positive definite matrix W such that (for all x and y in V) x * y = x'L'WLy. (c) Show that g is an inner product (for V) if and only if there exists an n x n symmetric positive definite matrix W such that (for all x and y in V) x * y = x'Wy. Solution, (a) Suppose that, for some /, x * y = (LxWLy) (for all x and y in V). Then, (1) x*y = (Lx)-(Ly) = (Ly)-(Lx)=y*x; (2) x * x = (Lx) • (Lx) > 0, with equality holding if and only if Lx = 0 or equivalently (since x = Bk for some vector k, so that Lx = 0 =$ LBk = 0 =*► Ik = 0 => k = 0 => Bk = 0 => x = 0) if and only if x = 0; (3) (Ax) * y = (ALx)-(Ly) = *[(Lx)-(Ly)] = *(x * y); (4) (x+y)*z = (Lx+Ly)-(Lz) = [(Lx)-(Lz)]+[(Ly)-(Lz)] = (x*z)+(y*z) (where x, y, and z represent arbitrary vectors in V and k represents an arbitrary scalar). Thus, g is an inner product. Conversely, suppose that g is an inner product, and consider the function / that assigns to an arbitrary pair of vectors s and t in lZrx l the value s*t=(Bs)*(Bt). We find that (1) s*t = (Bs)*(Bt) = (Bt)*(Bs) = Us; (2) s * s = (Bs) * (Bs) > 0, with equality holding if and only if Bs = 0 or equivalently (since the columns of B are linearly independent) if and only if s = 0; (3) (ks) • t = (*Bs) * (Bt) = fc[(Bs) * (Bt)] = *(s • t); (4) (s+t)*u = (Bs+Bt)*(Bu) = [(Bs)*(Bu)]+[(Bt)*(Bu)] = (s*u)+(t*u) (where s, t, and u represent arbitrary vectors in TZrx l and k represents an arbitrary scalar). Thus, / is an inner product (for TZrxl). Now, set / = /. Then, letting x and y represent arbitrary vectors in V and defining s and t to be the unique vectors that satisfy Bs = x and Bt = y (so that s = Is = LBs = Lx and similarly t = Ly), we find that x *y = (Bs) * (Bt) = s*t = s-t = (Lx)-(Ly). (b) Let / represent an arbitrary inner product for ftrxl, and denote by s*t the value assigned by / to an arbitrary pair of /-dimensional vectors s and t. According to Part (a), g is an inner product (for V) if and only if there exists an / such that (for all x and y in V) x * y = (Lx) • (Ly). Moreover, according to the discussion of
104 14. Linear, Bilinear, and Quadratic Forms Section 14.10a, every inner product for 7^rxl is expressible as a bilinear form, and a bilinear form (in r-dimensional vectors) qualifies as an inner product for TZrxl if and only if the matrix of the bilinear form is symmetric and positive definite. Thus, g is an inner product (for V) if and only if there exists anrxr symmetric positive definite matrix W such that (for all x and y in V) x * y = (Lx)'WLy. (c) Suppose that there exists an n x n symmetric positive definite matrix W such that (for all x and y in V) x * y = x'Wy. According to the discussion of Section 14.10a, the function that assigns the value x'Wy to an arbitrary pair of vectors x and y in TZ"xl is an inner product for %nxl. Thus, it follows from the discussion of Section 6.1b that g is an inner product (for V). Conversely, suppose that g is an inner product. According to Theorem 4.3.12, there exist n—r n-dimensional column vectors br+j bn such that b\ br, br+i b„ form a basis for TZnx l. Let C = (br+i, br+2 b„), and define F=(B, C). Partition F-1 as where L* is of dimensions r x n. (The matrix F is invertible since its columns are linearly independent.) Note that (by definition) L*B = Ir, so that L* is a left inverse of B, and MB = 0. According to Part (b), there exists an r x r symmetric positive definite matrix W* such that (for all x and y in V) x*y = x'L;cW*L*y. Moreover, letting x and y represent arbitrary vectors in V and defining s and t to be the unique vectors that satisfy Bs = x and Bt = y, we find that x'L;w*L*y = s'(L*B)'W+L*Bt = x'Wy, where W = (F_1 )/diag(W!„ I^-rJF-1. Thus, x * y = x'Wy. Furthermore, W is symmetric, and it follows from Lemma 14.8.3 and Corollary 14.2.10 that W is positive definite.
14. Linear, Bilinear, and Quadratic Forms 105 EXERCISE 42. Let V represent a linear space of m x n matrices, and let A*B represent the value assigned by a quasi-inner product to any pair of matrices A and B (in V). Show that the set W = {A6V:A«A = 0}, which comprises every matrix in V with a zero quasi norm, is a linear space. Solution. Let A and B represent arbitrary matrices in U% and let k represent an arbitrary scalar. Since (by definition) || A || = 0, it follows from the discussion in Section 14.10c that A-B = 0. Thus, (A + B)-(A + B) = (A-A) + 2(A-B) + (B-B) = 0 + 0 + 0 = 0, implying that (A + B) e U. Moreover, (*A)-(*A) = *2(A-A) = k2(0) = 0, so that A-A e U. We conclude that U is a linear space. EXERCISE 43. Let W represent an m x m symmetric positive definite matrix and V an n x n symmetric positive definite matrix. (a) Show that the function that assigns the value tr(A'WBV) to an arbitrary pair of m x n matrices A and B qualifies as an inner product for the linear space Hm x". (b) Show that the function that assigns the value tr(A'WB) to an arbitrary pair of m x n matrices A and B qualifies as an inner product for TZmx". (c) Show that the function that assigns the value tr(A'WBW) to an arbitrary pair of m x m matrices A and B qualifies as an inner product for lZmxm. Solution, (a) Let us show that the function that assigns the value tr(A'WBV) to an arbitrary pair of m x n matrices A and B has the four basic properties (described in Section 6.1b) of an inner product. For this purpose, let A, B, and C represent arbitrary m x n matrices, and let k represent an arbitrary scalar. (1) Using results (5.1.5) and (5.2.3), we find that tr(A'WBV) = tr[(A'WBV)'] = trCVB'WA) = tr(B'WAV). (2) According to Corollary 14.3.13, W = Q'Q for some m x m nonsingular matrix Q, and V = P'P for some n x n nonsingular matrix P. Thus, using results (5.2.3) and (5.2.5) along with Lemma 5.3.1, we find that tr(A'WAV) = trfA'Q'QAP'P) = trfPA'Q'QAP') = trKQAPYQAP'] > 0. with equality holding if and only if QAP; = 0 or, equivalently, if and only if A = 0.
106 14. Linear, Bilinear, and Quadratic Forms (3) Clearly, tr[(fcA)'WBV)] = k tr(A'WBV). (4) Clearly, tr[(A + B)'WCV] = tr(A'WCV) + tr(B'WCV). (b) and (c) The functions described in Parts (b) and (c) are special cases of the function described in Part (a) — those where V = I and V = W, respectively. EXERCISE 44. Let A represent a q x p matrix, B a p x n matrix, and C an m x q matrix. Show that (a) CAB(CAB)~C = C if and only if rank(CAB) = rank(C), and (b) B(CAB)~CAB = B if and only if rank(CAB) = rank(B). Solution, (a) Suppose that rank(CAB) = rank(C). Then, it follows from Corollary 4.4.7 that C(CAB) = C(C) and hence that C = CABR for some matrix R. Thus, CAB(CAB)~C = CAB(CAB)~CABR = CABR = C. Conversely, suppose that CAB(CAB)~C = C. Then, rank(CAB) > rank[CAB(CAB)~C] = rank(C). Since clearly rank(CAB) < rank(C), we have that rank(CAB) = rank(C). (b) Similarly, suppose that rank(CAB) = rank(B). Then, it follows from Corollary 4.4.7 that ft(CAB) = 7£(B) and hence that B = LCAB for some matrix L. Thus, B(CAB)~CAB = LCAB(CAB)~CAB = LCAB = B. Conversely, suppose that B(CAB)~CAB = B. Then, rank(CAB) > rank[B(CAB)~CAB] = rank(B). Since clearly rank(CAB) < rank(B), we have that rank(CAB) = rank(B). EXERCISE 45. Let U represent a subspace of K"xl, let X represent an n x p matrix whose columns span U, and let W and V represent nxn symmetric positive definite matrices. Show that each of the following two conditions is necessary and sufficient for the projection Px.wy of y on W with respect to W to be the same (for every y in 1Z") as the projection Px,vy of y on U with respect to V: (a) V = Px>wVPx.w + d- Px.w/Vd - Px.w); (b) there exist a scalar c, a p x p matrix K, and an n x n matrix H such that V = cW + WXKX'W + (1- Pxav/HU - Px.w). Solution, (a) In light of Theorem 14.12.18, it suffices to show that this condition is equivalent to the condition (I — Px.\v);VPx.w = 0. Suppose that V = Px wVPx.w + (I - PxAv/Vd - Px,w). Then, (l-Px.w/VPx.w = [Px.wd - Px.w)rVPX-W + [(I - Px.w/^Vd - Px.w)Px.w.
14. Linear, Bilinear, and Quadratic Forms 107 Moreover, since [according to Part (6) of Theorem 14.12.11] Px.w is idempotent, Px.wd - Pxw) = Pxw - Px,w = pxw - Pxw = 0, and similarly (I - Px.w)Px.w = 0. Thus, (I - Px.w)'VPx.w = 0. Conversely, suppose that (I - Px.w)'VPx,\y = 0. Then, V = [Px,w + (1- Px.wJl'VfPx.w + (1- Px.w)l = px.wVPx.w + d- Px.w)'V(I - Px.w) +(1 - Px.w)'VPx.w + [(I - Px.w)'VPx,w]' = Px.wVPx.w + d- Px.w)'V(I - Px.w). (b) Suppose that Px.wy is the same for every y in TZ" as Px.vy. Then, Condition (a) of the exercise is satisfied, so that V = PXWVPX w + (1- Px.w>'V(I - Px.w) = cW + WXKX'W +(1- Px.w)Hd - Px.w) for c = 0, K = [(X'WXn'X'VXfX'WX)-. and H = V. Conversely, suppose that V = cW + WXKX'W + (1- PX.W)'H(I - Px,w) for some scalar c and some matrices K and H. Then, since [according to Part (1) of Theorem 14.12.11] X - PX,WX = 0, we have that VX = cWX + WXKX'WX + (I - PX.W),H(X - PXtWX) = cWX + WXKX,WX = WX(cI + KX'WX) = WXQ for Q = cl + KX'WX. Thus, it follows from Theorem 14.12.18 that PXtWy is the same for every y in TZ" as Px.vy- EXERCISE 46. Let y represent an n -dimensional column vector, let U represent a subspace of 7£"xl, and let X represent an n x p matrix whose columns span U. "Recall" that, for any m x n matrix L, for any subspace W of 7£"xl, and for V = {v 6 TZm : v = Lx for some x e W}, x±HW <£> Lx±iV, (*) where H = LXL (and where x represents an arbitrary n x 1 vector), and C(LZ) = V, (**) where Z is any nxq matrix whose columns span W. Use results (*) and (**) to show that, for any nxn symmetric nonnegative definite matrix W, y ±w W if and onlyifX,Wy = 0.
108 14. Linear, Bilinear, and Quadratic Forms Solution. According to Corollary 14.3.8, W = L'L for some matrix L. Denote by m the number of rows in L, and let V = {v e TZm : v = Lx for some x e U). Then, it follows from result (*) [or equivalently from Part (3) of Lemma 14.12.2] that y ±w U if and only if Ly ±i V and hence, in light of result (**) [or equivalently in light of Part (5) of Lemma 14.12.2], if and only if (LX)'Ly = 0. Since (LX)'Ly = X'Wy, we conclude that y ±w U if and only if X'Wy = 0. EXERCISE 47. Let W represent an n x n symmetric nonnegative definite matrix. (a) Show that, for any n x p matrix X and any n x q matrix U such that C(U) C C(X), (1) WPx.wU = WU, and U'WPx.w = U'W; (2) Pu,wPx,w = Pu.w, and P^WP^w = WPx,wPu.w = WPU>W. (b) Show that, for any n xp matrixXand any n xq matrix Usuch thatC(U) = C(X), WPu,w = WPx.w. Solution, (a) (1) According to Lemma 4.2.2, there exists a matrix F such that U = XF. Thus, making use of Parts (1) and (5) of Theorem 14.12.25, we find that WPx.wU = WPX.WXF = WXF = WU and U'WPxw = F'X'WPx.w = FX'W = l/W. (2) Making use of Part (1), we find that Pu.wPx.w = U(U,WU)"U,WPx.w = UttfWUTU'W = Pu.w and similarly that WPx.wPu.w = WPx.wU(U,WU)-U,W = WUtfAVUTU'W = WPUW Further, making use of Part (3) of Theorem 14.12.25, we find that Px,wWPu.w = (WPxw/Puw = WPxwPu.w (b) Making use of Part (a), we find that WPu.w = W(Px.wPu.w) = WPxw- EXERCISE 48. Let U represent a subspace of 7£"xl, let A represent an n x n matrix and Wannx/i symmetric nonnegative definite matrix, and letX represent any n x p matrix whose columns span U.
14. Linear, Bilinear, and Quadratic Forms 109 (a) Show that A is a projection matrix for U with respect to W if and only if A = Px.w + (I-Px.w)XK for some p x n matrix K. (b) Show that if A is a projection matrix for U with respect to W, then WA = WPx.w. Solution, (a) Suppose that A = Px.w + (I-Px,w)XK for some matrix K. Then, for any w-dimensional column vector y, Ay = Px.wy + (1- Px,w)X(Ky). Thus, it follows from Corollary 14.12.27 that A is a projection matrix for U with respect to W. Conversely, suppose that A is a projection matrix for li with respect to W. Then, Ay e U for every y in ft", so that C(A) C U = C(X) and hence A = XF for some matrix F. Moreover, it follows from Parts (2) and (3) of Theorem 14.12.26 that WAy = WX(X'WX)~~X'Wy for every y in 11" [since one solution to linear system (12.4) is X'WXrX'Wy], implying that WA = WX(X,WX)"X,W and hence that WXF = WX(X'WX)-X'W. (S.8) Since [according to Part (1) of Theorem 14.12.25] (X'WX)-X' is a generalized inverse of WX, we conclude, on the basis of Theorem 11.2.4 and Part (5) of Theorem 14.12.25, that there exists a matrix K such that F = (X'WXrX'W + [I - (X'WX)-X'WX]K and hence such that A = X(X,WX)"X'W + X[I - (X,WX)"X/WX]K = Px.w + (1- Px.w)XK. (b) Suppose that A is a projection matrix for U with respect to W. Then, it follows from Part (a) that A = Px.w + (I-Px.w)XK for some matrix K. Thus, making use of Part (1) of Theorem 14.12.25, we find that WA = WPx.w + (WX - WPXWX)K = WPx,w
110 14. Linear, Bilinear, and Quadratic Forms EXERCISE 49. Let A represent an n x n matrix and W an n x n symmetric nonnegative definite matrix. (a) Show (by, e.g., using the results of Exercise 48) that if A'WA = WA [or, equivalently, if (I — A)'WA = 0], then A is a projection matrix with respect to W, and in particular A is a projection matrix for C(A) with respect to W, and, conversely, show that if A is a projection matrix with respect to W, then A'WA = WA. (b) Show that if A is a projection matrix with respect to W, then in particular A is a projection matrix for C(A) with respect to W. (c) Show that A is a projection matrix with respect to W if and only if WA is symmetric and WA2 = WA. Solution, (a) Suppose that A'WA = WA and hence that A'W = (WA)' = (A'WA/ = A'WA. Then, A = PA.W + A - A(A'WA)"A'W = PA>W + A - A(A'WA)"A'WA = PA.w + a-PA.w)AI,„ and it follows from Part (a) of Exercise 48 that A is a projection matrix for C(A) with respect to W. Conversely, suppose that A is a projection matrix with respect to W. Let U represent any subspace of TZ"X' for which A is a projection matrix with respect to W, and let X represent any n x p matrix whose columns span U. Then, according to Part (b) of Exercise 48, WA = WPXW, and, making use of Part (6 ) of Theorem 14.12.25, we find that A'WA = A'WPx.w = (WA)'Px,w = (WPx.w)'Px.w = PxwWPx.w = WPx.w = WA. (b) Suppose that A is a projection matrix with respect to W. Then, it follows from Part (a) that A'WA = WA, and we conclude [on the basis of Part (a)] that A is a projection matrix for C(A) with respect to W. (c) In light of Part (a), it suffices to show that A'WA = WA if and only if WA is symmetric and WA2 = WA. If WA is symmetric and WA2 = WA, then A'WA = (WA)'A = WAA = WA2 = WA. Conversely, if A'WA = WA, then (WA)' = (A'WA)' = A'WA = WA
14. Linear, Bilinear, and Quadratic Forms 111 (i.e., WA is symmetric), and WA2 = WAA = (WA)'A = A'WA = WA. EXERCISE 50. Let U represent a subspace of ft"xl, let X represent an n x p matrix whose columns span U, and let W and V represent /ixn symmetric nonnegative definite matrices. Show (by, e.g., making use of the result of Exercise 46) that each of the following two conditions is necessary and sufficient for every projection of y on U with respect to W to be a projection (for every y in 1Z") of y on U with respect to V: (a) X'VPx.w = X'V, or, equivalently, X'V(I - Px.w) = 0; (b) there exists a p x p matrix Q such that VX = WXQ, or, equivalently, C(VX) C C(WX). Solution, (a) It follows from the result of Exercise 46 that a vector z (in U) is a projection of a vector y (in 11") on U with respect to V if and only if X,V(y-z) = 0. Further, it follows from Corollary 14.12.27 that every projection of y on U with respect to W is a projection (for every y in 11") of y on U with respect to V if and only if, for every y and every vector-valued function k(y), X'Vfy - Px.wy - (I - Px.w)Xk(y)] = 0, or, equivalently, if and only if, for every y and every vector-valued function k(y), X'V(I - Px.w)[y - Xk(y)] = 0. (S.9) Thus, it suffices to show that condition (S.9) is satisfied for every y and every vector-valued function k(y) if and only if X'V(I — Px.w) = 0. If X'V(I — Px.w) = 0, then condition (S.9) is obviously satisfied. Conversely, suppose that condition (S.9) is satisfied for every y and every vector-valued function k(y). Then, since one choice for k(y) is k(y) = 0, x,va-Px.w)y = o for every y, implying that X'V (I - Px,w) = 0. (b) It suffices to show that Condition (b) is equivalent to Condition (a) or, equivalently [since VX = (X'V/ and PX.WVX = (X'VPx.w)']. to the condition vx = Px.wvx- <s-10> If condition (S.10) is satisfied, then VX = WX[(X'WXr]'X'VX = WXQ
112 14. Linear, Bilinear, and Quadratic Forms for Q = [(X/WX)-]'X/VX. Conversely, if VX = WXQ for some matrix Q, then, making use of Part (4) of Theorem 14.12.25, we find that px,wVX = PX,WWXQ = WXQ = VX, that is, condition (S.10) is satisfied. EXERCISE 51. Let X represent an n x p matrix and W an n x n symmetric nonnegative definite matrix. As in the special case where W is positive definite, let Cif(X) = iyennxl:y±v/C(X)}. (a) By, for example, making use of the result of Exercise 46, show that CW(X) = JV(X'W) = C(l - Px,w). (b) Show that dim[Cw(X)] =/i- rank(WX) > n - rank(X) = n - dim[C(X)]. (c) By, for example, making use of the result of Exercise 46, show that, for any solution b* to the linear system X'WXb = X'Wy (in b), the vector y — Xb* is a projection of y on C^(X) with respect to W. Solution, (a) It follows from the result of Exercise 46 that an /i-dimensional column vector y and C(X) are orthogonal with respect to W if and only if X'Wy = 0. Thus, C^(X) = J\f(X"W). Moreover, since [according to Part (5) of Theorem 14.12.25] X(X'WX)~ is a generalized inverse of X'W, we have (in light of Corollary 11.2.2) that ^(X'W) = C(l - Px.w). (b) Making use of Part (a), together with Part (10) of Theorem 14.12.25 and Corollary 4.4.5, we find that dim[Cw(X)] = dim[C(I - Px.w)] = rank(I - Px.w) = n - rank(WX) > n - rank(X) = n - dim[C(X)]. (c) Let z = Xb*. According to Theorem 14.12.26, z is a projection of y on C(X) with respect to W. Thus, (y - z) JLW C(X), and, consequently, (y - z) e Cw(X). It remains to show that [y - (y - z)] JLW CW(X) or, equivalent^, that z JLW CW(X). Making use of Part (4) of Theorem 14.12.25, we find that (I - Pjcw/Wz = (I - Px.w/WXb* = 0, implying (in light of the result of Exercise 46) that zJLwC(I-Px.w) and hence [in light of the result of Part (a)] that zJLCw(X).
15 Matrix Differentiation EXERCISE 1. Using the result of Part (c) of Exercise 6.2, verify that every neighborhood of a point x in TZmxl is an open set. Solution. Take the norm for TZmxl to be the usual norm, let N represent the neighborhood of x of radius r, and let y represent an arbitrary point in N. Further, take M to be the neighborhood of y of radius s=r-\\y-x\l and let z represent an arbitrary point in M. Then, using the result of Part (c) of Exercise 6.2, we find that || z - x || < || z - y || + || y - x || < s + || y - x || = r - || y - x || + || y - x || = r, implying that z € N. It follows that M C N and hence that y is an interior point of N. We conclude that TV is an open set. EXERCISE 2. Let / represent a function, defined on a set 5, of a vector x = (jci, ..., xm)f of m variables, suppose that the set S contains at least some interior points, and let c represent an arbitrary one of those points. Verify that if / is k times continuously differentiable at c, then / is k times continuously differentiable at every point in some neighborhood of c.
114 15. Matrix Differentiation Solution. Suppose that / is k times continuously differentiable at c. Then, there exists a neighborhood Nofc such that all of the first- through fcth-order partial derivatives of / exist and are continuous at every point in N. Let x* represent an arbitrary point in N. Since any neighborhood is an open set, x* is an interior point of N. Thus, there exists a neighborhood N* of x*, all of whose points belong to N. It follows that the first- through fcth-order partial derivatives of / exist and are continuous at every point in N* and hence that / is continuously differentiable at x*. EXERCISE 3. Let X = [x,j} represent an m x n matrix of mn variables, and letx represent an /wz-dimensional column vector obtained by rearranging the elements of X (in the form of a column vector). Further, let S represent a set of X-values, and let S* represent the corresponding set of x-values (i.e., the set obtained by rearranging the elements of each m x n matrix in S in the form of a column vector). Verify that an /wn-dimensional column vector is an interior point of S* if and only if it is a rearrangement of anmxn matrix that is an interior point of 5. Solution. Let C = {cyy} represent a value of X, and let c represent the corresponding value of x. Then (when the inner products for 7lm,lxl and fcmxn are taken to be the usual inner products) || X - C || = [(X - C)'(X - C)]1'2 = £ {Xij - Cijf = {tr[(X-C)'(X-C)]}1/2 = || X — C ||. Thus, a set of X-values is a neighborhood of C of radius r if and only if the corresponding set of x-values is a neighborhood of c of radius r. It follows that there exists a neighborhood of c, all of whose points belong to 5*, if and only if there exists a neighborhood of C, all of whose points belong to S. We conclude that c is an interior point of S* if and only if C is an interior point of S. EXERCISE 4. Let / represent a function whose domain is a set S in ft,nx l (that contains at least some interior points). Show that the Hessian matrix H/ of / is the gradient matrix of the gradient vector (D/)' of /. Solution. The gradient vector of / is (D/)' = (D\ /,..., Dmf)'. The gradient matrix of this vector is the m x m matrix whose ijth element is the /th (first-order) partial derivative Df.f of Dj /, which by definition is the Hessian matrix of /. EXERCISE 5. Let g represent a function, defined on a set 5, of a vector x = (a'i xmY ofm variables, let S* = [x e S : g(\) # 0}, and let c represent any interior point of S* at which g is continuously differentiable. Show that (for any
15. Matrix Differentiation 115 positive integer k) g k is continuously differentiable at c and dxj s dxj ' Do so based on the result that and, letting / (like g) represent a function (defined on S) of x, the results that if / is continuously differentiable at c, then the ratio f/g is also continuously differentiable at c, that, for any positive integer k, fk is continuously differentiable at any point at which / is continuously differentiable, and that dxj'1"1 Bxj' (**> Solution. In light of the results cited in the statement of the exercise (or equiva- lently in light of Lemma 15.2.2 and the ensuing discussion), we have that \/g is continuously differentiable at c and that, as a consequence, (1 /g)k or equivalently g~k is continuously differentiable at c. Moreover, using results (**) and (*) [or, equivalently, results (2.16) and (2.8)], we find that !p . mi* =^-^=,0/^-(-.)(1/^ dXj dXj aXj dxj EXERCISE 6. Let F represent a p x p matrix of functions, defined on a set 5, of a vector x = (x\,..., xm)' of m variables. Let c represent any interior point of S at which F is continuously differentiable. Show that if F is idempotent at all points in some neighborhood of c, then (at x = c) F^F = 0. absolution. Suppose that F is idempotent at all points in some neighborhood of c. Then, differentiating both sides of the equality F = FF [with the help of result (4.3)], we find that (at x = c) »,» + »,. «s„ dxj dxj dxj
116 15. Matrix Differentiation Premultiplying both sides of equality (S.l) by F gives , 3F „ 3F „ „ 3F 2—+F—F = F— „9F -3F ^aF^ 3F aF., F— = F2— + F—F = F— + F—F or equivalently f|^f = o. dxj EXERCISE 7. Let g represent a function, defined on a set 5, of a vector x = (jq,..., xm)' of m variables, and let f represent apxl vector of functions (defined on 5) of x. Let c represent any interior point (of 5) at which g and f are continuously differentiable. Show that gf is continuously differentiable at c and that (at x = c) B(gt) _ dg dt dx! dx! 8dx!' Solution. It follows from result (4.9) (and the discussion thereof) that gf is continuously differentiable at c and that (at x = c) d(gt) = dg df dxj dxj dxj 0' = 1 m). Moreover, since d(gf)/dxj, (dg/dxj)f, and g{df/dxj) are the yth columns of d(gf)/dx't f(dg/dx!\ and g(df/dx')t respectively, we have that £(£f) to _3f dx! dx! 8dx'' EXERCISE 8. (a) Let X = {*/;} represent an m x n matrix of mn "independent" variables, and suppose that X is free to range over all of 1Zmxn. (1) Show that, for any p xm and nx p matrices of constants A and B, atr(AXB) _ ., ax = A'B'. [Him. Observe that tr(AXB) = tr(BAX).] (2) Show that, for any m- and n-dimensional column vectors a and b, 3(a'Xb) -8X"=ab- [Hint. Observe that a'Xb = tr(a'Xb).] (b) Suppose now that X is a symmetric (but otherwise unrestricted) matrix (of dimensions m x m).
15. Matrix Differentiation 117 (1) Show that, for any p xm and m x p matrices of constants A and B, atr(AXB) ^ ——— = C + C - diag(cj j, c22 cmm\ where C = [cy] = BA. (2) Show that, for any wi-dimensional column vectors a = {at} and b = {bt}, dX = ab' + ba' - diagteifcj, a2b2 ambm). Solution, (a) (1) Using result (6.5), we find that atr(AXB) 3tr(BAX) ax ax = (BA)' = A'B'. (2) Observing that a'Xb = tr(a'Xb) and applying Part (1) (with A = a' and B = b), we find that a(a'Xb) ■ =ao . ax (b) (1) Using result (6.7), we find that atr(AXB) atr(BAX) ax ax C + C'-diag(cn, C22 cmm). (2) Observing that a'Xb = tr(a'Xb) and applying Part (1) (with A = a' and B = b), we find that d(a'Xb) 8X =ba +ab -diag(«i^i, a2b2 ambm). EXERCISE 9. (a) Let X = [xst) represent an m x n matrix of "independent" variables, and suppose that X is free to range over all offc"1*". Show that, for any nxm matrix of constants A, 3tr[(AX)2] ^^=2(AXA). (b) Let X = [xst} represent snm x m symmetric (but otherwise unrestricted) matrix of variables. Show that, for any m x m matrix of constants A, atrIltX)2] = 2(B + B' - diag(fc„, fc2. .... bmm)], where B = [bst) = AXA. Solution. Let uy- represent the j th column of an identity matrix (of unspecified dimensions).
118 15. Matrix Differentiation (a) According to results (4.7) and (5.3), 3(AXA) 3X . . , . —r = A-—A = Au/U.A. dxu dxu J Thus, making use of results (6.3). (5.3), and (5.2.3), we find that 3tr[(AX)2] 3tr[(AXA)X] dxij L„> 9X\ r 3(AXA)1 = tr( AXA-— ) + tr X-^- V BxijJ I dxu J = tr(AXAu,Uy) + tr(XAu/u}A) = 2tr(UyAXAu,-) = 2uyAXAuf. Since u'. AXAu,- is the jith element of AXA or, equivalently, the ijth element of (AXA)', we find that 9X (b) For purposes of differentiating tr[(AX)2], tr[(AX)2] is interpreted as a function of an 111(111 + 1 )/2-dimensional column vector x whose elements are xy (j </ = l m). According to results (4.7), (5.6), and (5.7), 8(AXA) _ _3X _ jAu/uJA, if j = i, dx.j ~ dxij ~|A(u/u;. + uyu;)A, if y</, and, according to result (6.3), 8tr[(AX)2] 8tr[(AXA)X] dxij dxij (\^l 3X\ r 3(AXA)1 = tr AXA-— ) + tr X-^- . Thus, making use of results (5.6), (5.7), and (5.2.3), we find that atr[(AX)-] = ^^j^^^j + tr(XAu/u;.A) oxa = 2tr(ul'AXAul) = 2 ujBu, and that (for j < i) atr[(AX)2] ihi; = tr[AXA(u,u} + uyuj)! + tr[XA(u/u} + uyii{)A]
15. Matrix Differentiation 119 = 2tr[AXA(u/U,;+u;u;.)] = 2 tr(AXAu,u';) + 2 tr(AXAuyu;.) = 2 tr(u'y AXAu,-) + 2 tr(uj AXAu;) = 2u^.Bi]f+2u{Buj. Since (for i,j = l m) u'Bu,- is the //th element of B or, equivalently, the ijth element of B' and since uJBuy is the ijth element of B, we conclude that 8tr[(^X)"] = 2 [B + B' - diag(*>,ub22 bmm)]. EXERCISE 10. Let X = {a,,} represent an m x n matrix of "independent" variables, and suppose that X is free to range over all of Hmxn. Show that, for Solution. Let u/ represent the yth column of I,,,. Then, making use of results (6.1), (4.8), (5.2.3), and (5.3), we find that 8tr(X*) _ /ax*\ Bxtj '^ydxij) x*-1-— + xk~2-— X + • • • + — x*-1) dx,j dxij dxij J ="(*'-'H) = ittr(X*"lU|uJ) = k trO^X*"1^) Moreover, u'.X*-1u/ equals the y/th element of X*~l or equivalently the ijth element of (X*"1)'. Since (X*"1)' = (X')*"1, we conclude that 9tr(X*)/3^,, which is the ijth element of 3tr(X*)/3X, equals the ijth element of fc(X')*-1 and hence that
120 15. Matrix Differentiation EXERCISE 11. Let X = [xst} represent an m x n matrix of "independent" variables, and suppose that X is free to range over all of 1Z",X". (a) Show that, for any m x m matrix of constants A, atrtx'AX) , —^-= (A + A)X. (b) Show that, for any n x n matrix of constants A, atr(XAX') 3X ■=X(A + A'). (c) Show (in the special case where n = m) that, for any m x m matrix of constants A, atr(XAX) 3X - = (AX)' + (XA)'. (d) Use the results of Parts (a)-(c) to devise simple formulas for atr(X'X)/aX, atr(XX')/aX, and (in the special case where n = m) dtr(X2)/dX. Solution. Let uy represent the yth column of an identity matrix (of unspecified dimensions). (a) Since tr(X'AX) = tr(AXX'), it follows from result (6.2) that atr(X' :;ax) rAa(xx')i d.\ij Thus, making use of results (4.3), (4.10), (5.3), and (5.2.3), we find that = tr(AXuyu;.) + tr(Auiu'jX') = tr(u;AXu;) + uVyX'Au/) = u{AXuy+uJX'Aui. Since ujAXuy is the ijth element of AX and since u'X'Au,- is the jith element of X'A or equivalently the ijth element of (X'A)', we conclude that ^A^=AX+(X'A)' = (A + A')X. (b) Since tr(XAX') = tr(AX'X), it follows from result (6.2) that atr(XAX' a. kax') _ r atx'xn ■V/y "trL BXU X
15. Matrix Differentiation 121 Thus, making use of results (4.3), (5.3), and (5.2.3), we find that = tr(AX'u;Uy) + tr(AuyuJX) = u,yAX/ul+u;XAu;. Since u'y.AX'ii/ is the jith element of AX' or equivalently the ijth element of (AX7)' and since uJXAuy is the ijth element of XA, we conclude that atr(XAX') ,AYV,YA v/Aj_ak — = (AX ) + XA = X(A + A'). d\ (c) Since tr(XAX) = tr(AXX), it follows from result (6.2) that 8tr(XAX) dx Thus, making use of results (4.3), (5.3), and (5.2.3), we find that !H™ tr(AX^)+fr(A^X) Bxu \ dxijj \ dxtj J = tr(AXu/ir}) + tr(Au/U;X) = UyAXU, + UyXAllj . Since iiy AXii,- is the jith element of AX or equivalently the ijth element of (AX)' and since UyXAu,- is the jith element of XA or equivalently the ijth element of (XA)', we conclude that ^ = (AX)' + (XA)<. (d) Upon setting A = I in the formulas from Parts (a)-(c), we find that atr(X'X) atr(XX') [XAX) rA3(xxn ax ax and (in the special case where n = m) atr(X2) ■=2X ax = 2X'. EXERCISE 12. Let X = {*//} represent an m x m matrix. Let f represent a function of X defined on a set S comprising some or all m x m symmetric
122 15. Matrix Differentiation matrices. Suppose that, for purposes of differentiation, f is to be interpreted as a function of the [m(m + l)/2]-dimensional column vector x whose elements are x/y (j < i = 1,..., m). Suppose further that there exists a function g, whose domain is a set T of not-necessarily-symmetric matrices that contains S as a proper subset, such that g(X) = /(X) for X e S> so that g is a function of X and / is the function obtained by restricting the domain of g to S. Define S* = {x : X € S}. Let c represent an interior point of 5*, and let C represent the corresponding value of X. Show that if C is an interior point of T and if g is continuously differentiable at C, then f is continuously differentiable at c and that (at x = c) BX BX \BXJ B\Bxu Bx22 Bxmm) Solution. Let H represent the m x m matrix of functions defined, on S*, by H(x) = X. Then, H is continuously differentiable at c. Thus, it follows from the results of Section 15.7 that if C is an interior point of T and if g is continuously differentiable at C, then f is continuously differentiable at c and (at x = c) Bxij l\BXjBxij] U <i = 1 ml Moreover, in light of results (5.6), (5.7), and (5.2.3), we have that «[(S)'SH(IM-<I)« and that (for/ < i) -1(1)¾] «K4H«[(SM -«(»'-f(SK Since uj (9g/dX)'u; is the ith diagonal element of {Bg/BX)' (or equivalently the /th diagonal element of Bg/BX) and since Uy(dg/dX)'iif and uJ(3g/3X)'uy are the //th elements of Bg/BX and (Bg/BX)', respectively, it follows that (at x = c) V = ag_ (Bg\'_ . (Bg_ Bg_ Bg \ BX BX \BX) 8Uvii ' 3.V22 BxmmJ EXERCISE 13. Let h = [hi] represent an n x 1 vector of functions, defined on a set 5, of a vector x = (.\| .v,„)' of m variables. Let g represent a function, defined on a set 7\ of a vector y = (vi y„)' of n variables. Suppose that h(x) e T for every x in 5, and take f to be the composite function defined (on
15. Matrix Differentiation 123 S) by f(x) = g[h(x)]. Show that if h is twice continuously differentiable at an interior point c of S and if [assuming that h(c) is an interior point of T] g is twice continuously differentiable at h(c), then / is twice continuously differentiable at cand H/(c) = [Dh(c)]'Hg[h(c)]Dh(c) + Y, Dig[h(c)]Hhi(c). i=i Solution. Suppose that h is twice continuously differentiable at c or equivalently that h i hn are twice continuously differentiable at c. Suppose further that g is twice continuously differentiable at h(c). Then, g is continuously differentiable at h(c) and hence is continuously differentiable at every point in some neighborhood Ng of h(c). Moreover, hi hn are continuously differentiable at c and (in light of Lemma 15.1.1) continuous at c. Consequently, there exists a neighborhood Afy, of c such that (/) hi hn are continuously differentiable at every point in N/, and (ii) h(x) e Ng for every x in Thus, it follows from Theorem 15.7.1 that / is continuously differentiable at every x in N/, and that (for x e Nh) n Dy/(x) = ^M/(X)Dy/l/(x), /=1 where w,(x) = D/g[h(x)]. Since Dig is continuously differentiable at /i(c), we have (as a further consequence of Theorem 15.7.1) that Dsui(e) = Y/D2kig[h(c)]Dshk(c). Jt=i Since Djhi is continuously differentiable at c, we conclude that Djf (like f) is continuously differentiable at c and hence that f is twice continuously differentiable at c. Moreover, p*./(c) = DsDjf(c) n = J2 [ui(c)D*jh,{c) + DsUi(c)Djhi(e)] /=1 = £ DtgfhiOlDijhiic) + E E Dlg[h(c)}Dshk(c)Djhi(c) /=1 A-=l i=l = J2 Dig[h(c)]D*jhi{c) + [Dsh(c)]'Hg[h(c)]Djh(c). i=l To complete the argument, observe that D*jhj(c) is the sj\h element of H/i,(c) and that [£>5h(c)]'Hg[h(c)]D;h(c) is the sjth element of [Dh(c)]/Hg[h(c)]Dh(c)
124 15. Matrix Differentiation and hence that H/(c) = [Dh(c)]'Hg[h(c)]Dh(c) + £ Dig[h(c)}Wii(c). EXERCISE 14. Let X = [xjj) represent an m x m matrix of m2 "independent" variables (where m > 2), and suppose that the range of X comprises all ofR,mxm. Show that (for any positive integer k) the function f defined (on 1Zm xm) by /(X) = |X|* is continuously differentiable at every X and that ^=*|X|*-'[adj(X)]'. Solution. For purposes of differentiation, rearrange the elements of X in the form of an /ir-dimensional column vector x, and reinterpret / as a function of x (in which case the domain of f comprises all of Hm~). Let h represent a function of x defined (on Rm~) by /i(x) = det(X), let g represent a function of a variable v defined (for all v) by g(y) = v*, and express f as the composite of g and /i, so that fix) = g[h(x)l The function g is continuously differentiable at every v, and 8y ' And, the function h is continuously differentiable at every x, and where f/y is the cofactor of the z'/th element x-,j of X. Thus, it follows from the chain rule that f is continuously differentiable at every x (or equivalently at every X) and that 9/ i.ivi*-K or equivalently that d.\ij ■±=k\xri[*&}<x)Y. EXERCISE 15. Let F = {fis) represent a p x p matrix of functions, defined on a set 5, of a vector x = {x\ xmY of m variables. Let c represent any interior point (of S) at which F is continuously differentiable. Use the results of Exercise 13.10 to show that (a) if rank[F(c)l = /?-!, then (at x = c) 3det(F) , ,9F dXj ' BXj
15. Matrix Differentiation 125 where z = [zs] and y = {v/} are any nonnull p-dimensional vectors such that F(c)z = 0 and [F(c)]'y = 0 and where [letting fas represent the cofactor of fisic)] k is a scalar that is expressible as k = fas/iyiZs) for any i and s such that y/ # 0 and zs # 0; and (b) if rank[F(c)] < p - 2, then (at x = c) 3det(F) Bxj Solution. Recall that det(F) is continuously differentiable at c and that (at x = c) = tr adj(F)— . 8det(F) 3. (a) Suppose that rank[F(c)] = p - 1. Then, according to the result of Part (a) of Exercise 13.10, adj[F(c)] = *zy', so that (at x = c) st(F) / , 3F \ , /,3F\ _ , 3F 3det(F) 3. (b) Suppose that rank[F(c)] < p - 2. Then, it follows from the result of Part (b) of Exercise 13.10 that (at x = c) 3det(F) Bxj V BxjJ EXERCISE 16. Let X = [xst} represent an m x n matrix of "independent" variables, let A represent an m x m matrix of constants, and suppose that the range of X is a set S comprising some or all X-values for which det(X'AX) > 0. Show that log det(X'AX) is continuously differentiable at any interior point C of S and that (at X = C) aiogdetM) = AX(X,AX)_, + [(X,AX)-.x,Ar. 9X Solution. For purposes of differentiating a function of X, rearrange the elements of X in the form of an m«-dimensional column vector x and reinterpret the function as a function of x, in which case the domain of the function is the set S* obtained by rearranging the elements of each m x « matrix in S in the form of a column vector. Let c represent the value of x corresponding to the interior point C of S (and note that c is an interior point of S*). Since X is continuously differentiable at c, X'AX is continuously differentiable at c, and hence log det(X'AX) is continuously differentiable at c (or equivalently at C).
126 15. Matrix Differentiation Moreover, making use of results (8.6), (4.6), (4.10)» (5.3), and (5.2.3) and letting uy represent the jth column of lm or I/,, we find that (at x = c) aiogdet(x'AX) = r .agAX)] Bxtj L Bxtj J = trKX'AXr'X'Au/u';] +tr[(X/AX)-1UyU;.AX] = u^X'AXr'X'Au/ +u<AX(X'AXr1ui. Upon observing that u;AX(X/AX)_1uy and u^X'AX^X'Au/ are the i>th elements of AX(X'AX)-1 and [(X'AXr'X'A]', respectively, we conclude that (at x = c) 31ogdet(X'AX) = AX(X,AX)_, + [(x^-ix^y. 9X EXERCISE 17. (a) Let X represent an m x n matrix of "independent" variables, let A and B represent q x m and nxq matrices of constants, and suppose that the range of X is a set S comprising some or all X-values for which det(AXB) > 0. Show that logdet(AXB) is continuously differentiable at any interior point C of S and that (at X = C) 8logdet(AXB)=[B(AXB)_,A]( 3X (b) Suppose now that X is an m x m symmetric matrix; that A and B are q x m and m x q matrices of constants; that, for purposes of differentiating any function of X, the function is to be interpreted as a function of the column vector x whose elements are x,j (j < i = 1 w); and that the range of x is a set S comprising some or all x-values for which det(AXB) > 0. Show that log det(AXB) is continuously differentiable at any interior point c (of S) and that (at x = c) aiogdet(AXB) „ . jm „ dX = K + K' - diagfti, *22 A*,), where K = [ku} = B(AXB)"1 A. Solution, (a) For purposes of differentiating a function of X, rearrange the elements of X in the form of an w/z-dimensional column vector x and reinterpret the function as a function of x, in which case the domain of the function is the set S* obtained by rearranging the elements of each m x n matrix in S in the form of a column vector. Let c represent the value of x corresponding to the interior point C of S (and note that c is an interior point of S*). Then, X is continuously differentiable at c, implying that AXB is continuously differentiable at c and hence that log det(AXB) is continuously differentiable at c (or equivalently at C).
15. Matrix Differentiation 127 Moreover, in light of results (8.6), (4.7), (5.3), and (5.2.3), we have that (at x = c) (AXB^A^B OXij J 3logdet(AXB) T ^TOI dxu I dxu J = trl (AXBJ-^u/u^B = u'^AXBr^u/ and hence {since UyB(AXB)_1Au/ is the ijth element of [B(AXB)"1 A]'} that (at 31ogdet(AXB)_m/AVo , ax • = ^(AXB^A]'. (b) By employing essentially the same reasoning as in Part (a), it can be established that logdet(AXB) is continuously differentiable at the interior point c and that (at x = c) aiogdet(AXB) ^?>=J(AXB)-A^Bl. XU I dXij J Bxu Moreover, in light of results (5.6), (5.7), and (5.2.3), we have that tr| (AXB)"1 A^-B 1 = trKAXB^AuiujB] = ujKu, and that (for j <i) trj (AXB^A^-B 1 = tif(AXB)"'Au,ur}B] + trKAXB)"1 AuyujB] = UyKu/ + ujKuy. Since u-Ku,- is the ith diagonal element of K and since uJKuy and u'.Ku,- are the ijth elements of K and K', respectively, it follows that (at x = c) 81ogdet(AXB) A- ,u , — =K + K -diag(fcn, k22 kqq). EXERCISE 18. Let F = [fa} represent apxp matrix of functions, defined on a set 5, of a vector x = (*i xm)' of m variables, and let A and B represent q x p and p x q matrices of constants. Suppose that S is the set of all x-values for which F(x) is nonsingular and det[AF_1(x)B] > 0 or is a subset of that set. Show that if F is continuously differentiable at an interior point c of S, then log det(AF_1B) is continuously differentiable at c and (at x = c) 81ogdet(AF dx ■j L sxj J
128 15. Matrix Differentiation Solution. Suppose that F is continuously differentiable at c. Then, in light of the results of Section 15.8, AF_1B is continuously differentiable at c and hence log det(AF-1B) is continuously differentiable ate. Moreover, malcing use of results (8.6), (8.18), and (5.2.3), we find that (at x = c) aiogdeKAF-'B) = tr (AF-'B) , ,3(AF AF^B)! dxj J r #f n = tr (AF-'Bj-^-AF-1-—F-'B) L BxJ J = -trrF-lB(AF-|B)-1AF-1^-l. EXERCISE 19. Let A and B represent q x m and m x q matrices of constants. (a) Let X represent an m x m matrix of m2 "independent" variables, and suppose that the range of X is a set S comprising some or all X-values for which X is nonsingular and det(AX_1B) > 0. Use the result of Exercise 18 to show that log det(AX_1B) is continuously differentiable at any interior point C of S and that (atX = C) aiogdeKAX-'B) 3X • = -[X-^AX-'Br'AX-1]'. (b) Suppose now that X is an m x m symmetric matrix; that, for purposes of differentiating any function of X, the function is to be interpreted as a function of the column vector x whose elements are .*,-, (j < i = 1 m)\ and that the range of x is a set S comprising some or all x-values for which X is nonsingular and det(AX_1B) > 0. Use the result of Exercise 18 to show that log det(AX_1B) is continuously differentiable at any interior point c of S and that (at x = c) aiogdeKAX^B) Tr Tr/ ,. , , * ax ~ = -K - K' + diag(*u. *22 kqq), where K = {ku} = X~1B(AX-1B)-,AX-1. Solution, (a) For purposes of differentiating a function of X, rearrange the elements of X in the form of an /n2-dimensional column vector x and reinterpret the function as a function of x, in which case the domain of the function is the set S* obtained by rearranging the elements of each m x m matrix in S in the form of a column vector. Let c represent the value of x corresponding to the interior point C of S (and note that c is an interior point of S*). Since X is continuously differentiable at c, it follows from the result of Exercise 18 that logdet(AX_1B) is continuously differentiable at c and that (at x = c) 3.ogdet(AX-'B) = _Jr.b^b,-.^. « 1
15. Matrix Differentiation 129 Moreover, in light of results (5.3) and (5.2.3), we have that trFx-^CAX-^)-1 AX"1 ^-1 = tr[X-lB(AX-,B)-lAX-,ulu,>] = u^X"IB(AX"lB)"lAX-|Ui. And, upon observing that u^X~1B(AX~1B)-1AX~1u/ is the ijth element of [X-'BCAX-'B)"1 AX"1]', we conclude that a.ogdet(AX->B) ,-,,^^,^,, ax = -{X-lB(AX-lB)-lAX-1]'. (b) By employing essentially the same reasoning as in Part (a), it can be established that logdet(AX_1B) is continuously differentiable at the interior point c and that (at x = c) 81ogdet(AX-lB) 3* let(AX"lB) / 8X\ dxij ~ \ dxijj' Moreover, in light of results (5.6), (5.7), and (5.2.3), we have that tr(K^)=tr(KU,^) = U''KU/ and that (for j < i) \i(k^—\ = tr(Ku/u'y) + tr(Ku7u;) = u'jKui-{-u-Kuj. Since ufKu, is the /th diagonal element of K and since ujKuy and u'Ku,- are the ijth elements of K and K', respectively, it follows that (at x = c) aiogdeKAX-'B) __ v, t .. n , .„ = -K - K + diag(*n,*22 kqq). oX EXERCISE 20. Let F = [fjs) represent a p x p matrix of functions, defined on a set 5, of a vector x = (,vi *„,)' of m variables. Let c represent any interior point (of S) at which F is continuously differentiable. By, for instance, using the result of Part (b) of Exercise 13.10, show that if rank[F(c)] < p - 3, then aadj(F) = Q absolution. Let fcj represent the cofactor of fsi and hence the /\sth element of adj(F), and let F5t represent the (p - 1) x (p - 1) submatrix of F obtained by striking
130 15. Matrix Differentiation out the sih row and the ith column (of F). Then, as discussed in Section 15.8, ¢,,- is continuously differentiable at c and (at x = c) ^ = (-,)-^)¾. Now, suppose that rank[F(c)] < p - 3. Then, rank[F5,(c)] < p - 3 [since otherwise F5,(c), and hence F(c), would contain anrxr nonsingular submatrix, where /* > p — 3, in which case the rank of F(c) would exceed p — 3]. Thus, it follows from Part (b) of Exercise 13.10 that adjPMc)] = 0, leading to the conclusion that (at x = c) dfei/dxj = 0 and hence that (at x = c) 3adj(F) = 0 dxj EXERCISE 21. (a) Let X represent an m x m matrix of m2 "independent" variables, and suppose that the range of X is a set S comprising some or all X-values for which X is nonsingular. Show that (when the elements of X-1 are regarded as functions of X) X-1 is continuously differentiable at any interior point C of S and that (at X = C) ax-1 , 1^ = -^- where y; represents the ith column and z'. the yth row of X-1. (b) Suppose now that X is an m x m symmetric matrix; that, for purposes of differentiating a function of X, the function is to be interpreted as a function of the column vector x whose elements are xy (j < i = 1 m); and that the range of x is a set S comprising some or all x-values for which X is nonsingular. Show that X-1 is continuously differentiable at any interior point c of S and that (at x = c) ax^f-y,y;- if y = /. Bxu l-y/y'y-yyy,'. if ;' </ (where y,- represents the /th column of X-1). Solution. Denote by uy the jth column of 1,,,. (a) For purposes of differentiating a function of X, rearrange the elements of X in the form of an /n2-dimensional column vector x and regard the function as a function of x, in which case the domain of the function is the set obtained by rearranging the elements of each m x m matrix in S in the form of a column vector. Let c represent the value of x corresponding to the interior point C of S. Then, X is continuously differentiable at c, implying that X-1 is continuously differentiable at c and [in light of results (8.15) and (5.3)] that (at x = c) dX~ - y-l 8XY-I __*-!„.„'v-l V7> —— - -A —- A - -A U/UyA - -y,Z ■ .
15. Matrix Differentiation 131 (b) The matrix X is continuously differentiable at the interior point c, implying that X"1 is continuously differentiable at c and [in light of results (8.15), (5.6), and (5.7)] that (at x = c) ^=-^-=-^^=^ and similarly (for j < i) 3X_1 — = -X-'fti/u} + U;u;.)X-' = -ytfj - y,.y!. EXERCISE 22. Let X represent an m x m matrix of m1 "independent" variables. Suppose that the range of X is a set S comprising some or all X-values for which X is nonsingular, and let C represent an interior point of S. Denote the ijth element of X-1 by yijt the yth column of X-1 by yy-, and the ith row of X-1 by zj. (a) Show that X~! is twice continuously differentiable at C and that (at X = C) 32X_1 (b) Suppose that det(X) > 0 for every X in S. Show that logdet(X) is twice continuously differentiable at C and that (at X = C) 82logdet(X) dxijdxst Solution. For purposes of differentiating a function of X, rearrange the elements of X in the form of an m2-dimensional column vector x and reinterpret the function as a function of X, in which case the domain of the function is the set S* obtained by rearranging the elements of each m x m matrix in S in the form of a column vector. Let c represent the value of x corresponding to the interior point C of S (and note that c is an interior point of S*). Denote by uy- the j\h column of Im. It follows from the results of Section 15.5 (together with Lemma 15.4.1) that X is twice continuously differentiable at c and that (at x = c) dX/dxjj = u/u^ and d2X/dxijdx5t = 0. (a) Based on the results of Section 15.9, we conclude that X-1 is twice continuously differentiable at c and that (at x = c) —r— = x-Wx-Vu;x-1 +x-|u,ulx-W-X"1 dxjjOXst = yjsyrt+ytiys'*<'j.
132 15. Matrix Differentiation (b) Similarly, based on the results of Section 15.9 (along with Lemma 5.2.1), we conclude that logdet(X) is twice continuously differentiable at c and that (at x = c) 82fgf(X) = -tr(x-Vu;.x-Vu;) = -u;x-'u,.u}x-V dxijdxsr J J = -ynyjs • EXERCISE 23. Let F = [fiS] represent apxp matrix of functions, defined on a set S, of a vector x = (jrj xm)' of m variables. For any nonempty set T = [t\ ts}. whose members are integers between 1 and m, inclusive, define D(D = d5F/dxfl • • • dx,s. Let k represent a positive integer and, for i = 1 k, let jj represent an arbitrary integer between 1 and /h, inclusive. (a) Suppose that F is nonsingular for every x in S, and denote by c any interior point (of S) at which F is k times continuously differentiable. Show that F_1 is k times continuously differentiable at c and that (at x = c) 3*F-1 k — = £ £ (-l)rF-,D(7'1)F-ID(72)--F-1D(7;)F-1, (E.1) °xji~-°xjk ,=i r, Tr where T\ Tr are r nonempty mutually exclusive and exhaustive subsets of [j\ j/;} (and where the second summation is over all possible choices for 7-1 Tr). (b) Suppose that det(F) > 0 for every x in 5, and denote by c any interior point (of S) at which F is k times continuously differentiable. Show that logdet(F) is k times continuously differentiable at c and that (at x = c) 8*'logdet(F) dxj{ ■ • • dxjk k = £ Y^ (-Or+Itr[F-1D(7,i)F-ID(7,2)-F-,D(7;)], (E.2) r=I T\ Tr where T\ Tr are r nonempty mutually exclusive and exhaustive subsets of [j\ jk] with jk 6 Tr (and where the second summation is over all possible choices for T\ 7». Solution, (a) The proof is by mathematical induction. For k = 1 and k = 2, it follows from the results of Sections 15.8 and 15.9 that F~' is k times continuously differentiable and formula (E.1) valid at any interior point at which F is k times continuously differentiable. Suppose now that, for an arbitrary value of A\ F-1 is k times continuously differentiable and formula (E.1) valid at any interior point at which F is k times continuously differentiable. Denote by c* an interior point at which F is k + 1 times continuously differentiable. Then, it suffices to show that F~' is k + 1 times
15. Matrix Differentiation 133 continuously differentiable at c* and that (at x = c*) dxjt • • • dxjk+l k+\ = J2 Jl (-l),'F-1D(7,I*)F-,D(7,2*).••F-1D(7;*)F-,, (S.2) r=lT* T; where y*+i is an integer between 1 and /», inclusive, and where T{ T* are r nonempty mutually exclusive and exhaustive subsets of {j\ jk+i). The matrix F is k times continuously differentiable at c* and hence at every point in some neighborhood Nofc*. By supposition, F_l is k times continuously differentiable and formula (E.l) valid at every point in N. Moreover, all partial derivatives of F of order less than or equal to k are continuously differentiable at c*. Thus, it follows from results (4.8) and (8.15) that dkF-l/dxjt • ■ • dxjk is continuously differentiable at c* and that (at x = c*) a*+iF-i a.vy, • • • dxjk+l =£ E <-'>' /•=1 iT, Tr r #f x -F-1- F-'DmjF-'D^) L fajit+i -F-lD(Ti)F-l-^—F-lD(T2) dxjk+i -F-1D(7,j)F-1D(7,2) • • • F-1D(7»F"1 ^^F-1 +F-lD(TlU[jk+l))F-lD(T2)--'F-lD(Tr)F-1 +F-1D(7i )F-lD(T2 U [jk+l}) • • • rlD(rr)r' +F-1D(7,)F-,D(72) • ~F-lD(Tr U U+i})F_1l. (S.3) The terms of sum (S.3) can be put into one-to-one correspondence with the terms of sum (S.2) (in such a way that the corresponding terms are identical), so that formula (S.2) is valid and the mathematical induction argument is complete. (b) The proof is by mathematical induction. For k = 1 and k = 2, it follows from the results of Sections 15.8 and 15.9 that logdet(F) is k times continuously differentiable and formula (E.2) valid at any interior point at which F is k times continuously differentiable. •F^DO^F-1 ■F^Dd^F-1
134 15. Matrix Differentiation Suppose now that, for an arbitrary value of fc, log det(F) is k times continuously differentiable and formula (E.2) valid at any interior point at which F is k times continuously differentiable. Denote by c* an interior point at which F is k 4- 1 times continuously differentiable. Then, it suffices to show that log det(F) is k 4-1 times continuously differentiable at c* and that (at x = c*) 3*logdet(F) = J2 H (-l)r+1tr[F-1D(71*)F-1D(7,2*)-^-^(7^)1, (S.4) r=l T* T* where 7\* T* are r nonempty mutually exclusive and exhaustive subsets of Ui yjt+i} with yjt+i e 7>*. The matrix F is k times continuously differentiable at c* and hence at every point in some neighborhood N ofc*. By supposition, log det(F) is k times continuously differentiable and formula (E.2) valid at every point in N. Moreover, all partial derivatives of F of order less than or equal to k are continuously differentiable at c*, and F"1 is continuously differentiable at c*. Thus, it follows from results (4.8) and (8.15) that dk log det(F)/3jCy,... dxjk is continuously differentiable at c* and that (at x = c*) 9*+1logdet(F) k /-=17-, Tr {- -F-1 -^-F-1D(7,1)F-1D(7,2) • • • F~1D(7» 3F -F_1D(7,i)F-1 F_1D(72) • • • F^1D(7» dxh+i op -F-1D(7,i)F-1D(7,2) • • • F"1 F~lD(Tr) dxjk+i +F-1D(7, U {^+1^-^(72)-.^-^(7,) +F-1D(7-i)F-1D(72 U [jk+l]) • • • F-1D(7» +F-1D(7-i)F-1D(72) • • .F-'DO-r U 1/a+i})]. (S.5) The terms of sum (S.5) can be put into one-to-one correspondence with the terms of sum (S.4) (in such a way that the corresponding terms are identical), so that formula (S.4) is valid and the mathematical induction argument is complete.
15. Matrix Differentiation 135 EXERCISE 24. Let X = {jr/y} represent an m x m symmetric matrix, and let x represent the m(m + l)/2-dimensional column vector whose elements are jr/y (j < i = 1 m).DefineStobethesetofallx-valuesforwhichXisnonsingular and S* to be the set of all x-values for which X is positive definite. Show that S and S* are both open sets. Solution. Let c represent an arbitrary point in S, and c* an arbitrary point in S*. It suffices to show that c and c* are interior points (of S and 5*, respectively). Denote by C and C* the values of X at x = c and x = c*, respectively. According to Lemma 15.10.2, there exists a neighborhood N of c such that X is nonsingular for x e N. And, it follows from the very definition of S that N C 5. Thus, c is an interior point of 5. Now, let Xjt and Cjj! represent the fcth-order leading principal submatrices of X and C*, respectively. Then, det(Xjt) is a continuous function of x (at all points in ftm<",+1>/2) and hence lim det(Xjt) = det(Q). x-*c* Since (according to Theorem 14.9.5) det(Q) > 0, there exists a neighborhood N* of c* such that | det(Xjt) - det(Cp| < det(Q) for x e N* and hence {since -[det(Xjt) - det(Cp] < | det(Xjt) - det(Q)|} such that - det(Xjt) + det(Q) < det(Q)forx 6 N*.Thus,-det(Xjt) < Oforx e 7V*or,equivalently,det(Xjt) > 0 for x e N* (k = 1 m). Based on Theorem 14.9.5, we conclude that X is positive definite for x 6 N*, or equivalently that N* C 5*, and hence that c* is an interior point of S*. EXERCISE 25. Let X represent an n x p matrix of constants, and let W represent an n x n symmetric positive definite matrix whose elements are functions, defined on a set 5, of a vector z = {z\ zmY of m variables. Further, let c represent any interior point (of S) at which W is twice continuously differentiable. Show that W — WPx,w is twice continuously differentiable at c and that (at z = c) 32(W - WPx,w) BziBzj 32W 'dZiSzj* aw ,aw - (I-Pxw)T-X(X/WX)-X,—(I-Px.w) azi oZj aw 3W - [(I-Px.w)t-X(X'WX)-X'—(1-Px.w)]'- OZi OZj Solution. Since W is twice continuously differentiable at c, it is continuously differentiable at c and hence continuously differentiable at every point in some neighborhood N of c. Then, it follows from Theorem 15.11.1 that W - WPx,w is = (I-Px>w)^-^-(I-Px.w)
136 15. Matrix Differentiation continuously differentiable at every point in N and that (for z e N) 3(W-WPx,w) „ p' ,3W , bTj = (I" Pxw)ai7(I" Px'w)" Further, Px.w and 3W/3zy are continuously differentiable at c. Thus, 3(W — WPx.vf)/Bzj is continuously differentiable at c, and hence W — WPx.w is twice continuously differentiable at c. Moreover, making use of results (4.6) and (11.1) [along with Part (3') of Theorem 14.12.11], we find that (at z = c) a2(W-WPx,w) = 3[3(W-WPx,w)/3zy] dz = -(I-Px.w) BZidZj dZi 8W3Px.w Bzj Bzi 32W +<,-p»,fe<5;<,-p"'"> -K1f)s*H' 32W + (I-Px.w)/^^-(I-Px.w) dZiOZj 3W 3W = -[(I - Px w)_X(X'WX)-X'—(I - Px.w)]' OZi oZj 32W dzidZj* + (i-pXtW)7^^r-a-px.w) / 8W 3W - a - px,w)—x(x/wx)-x/—(i - px,w). OZi OZj EXERCISE 26. Let X represent an n x p matrix and W an n x n symmetric positive definite matrix, and suppose that the elements of X and W are functions, defined on a set 5, of a vector z = (z\ zm)' of w variables. And, let c represent any interior point (of S) at which W and X are continuously differentiable, and suppose that X has constant rank on some neighborhood of c. Further, let B represent any pxn matrix such that X'WXB = X'W. Then, at z = c, a(WPx.w) = 3w_(i_p. }aw oZj ozj ozj + W(I - Px.w)t-B + [W(I - PXtW)—B]'. (*) Bzj dzj
15. Matrix Differentiation 137 Derive result (*) by using the result px.wwpx.w = WPx.w (**) to obtain the representation 3(WPx.w) Pxw) p> w9Px.w,p' aw ,Ppx.w\' p — = px.wW— + px.w^px.w + (^-^- j WPx.w, n making use of the result (^T^) WPx.w = (I - Px w)^Px.w + W(I - Px.w)|^B. (*) \ oZj / ' ozj azj and by then making use of the result Solution. According to result (**) [or, equivalently, according to Part (6') of Theorem 14.12.11], WPx.w = PxwWPx.w. Thus, it follows from results (4.6) and 4.10) that 3(WPx.w) P' wapx.w _, aw ^/arx.wX' Substituting from result (•) [or equivalently from result (11.16)], we find that (for any p x n matrix B such that X'WXB = X'W) 3(WPx.w) rrI P' aw ax — = [(1 - Px w)^-px.w + W(l - Px.w)t—B] oZj OZj OZj , aw , aw + Pxw—-Px.w + (1- px w)_px.w dzj ™ * A'w,aZy ax„ -px.w)T-B dzj aw ¥ _ „_,_ _ ax ^, + W(I-Px.w)|^B dzj = px W7- (I - px.w) + [W(I - Px.w)—B]; * OZj oZj + ^px,w + W(I - Pk.w)|^B. dZj dZj And, upon reexpressing (3W/9z/)px.w as aw„ aw aw/¥ n x t—Px.w = -z t— (I - px.w), oZj oZj oZj it is clear that a(WPx,w) aw , aw —r = « (i - px.w)-^r (i - px.w) ozj ozj o^j QV 5V + W(I - Px.w)t-B + [W(I - Px.w)7-Br. OZj OZj
16 Kronecker Products and the Vec and Vech Operators EXERCISE 1. (a) Verify that, for any m x n matrices A and B and any p x q matrices C and D, (A + B)®(C + D) = (A®C) + (A®D) + (B®C) + (B®D). (b) Verify that, for any m x n matrices Ai, A2 Ar and p x q matrices Bi,B2 B„ (EA') ® (Z>) = EE (Ai ®By). Solution, (a) It follows from results (1.11) and (1.12) that (A + B)®(C + D) = [A®(C + D)] + [B®(C + D)] = (A®C) + (A®D) + (B®C) + (B®D). (b) Let us begin by showing that, for any m x n matrix A, A® nTBy ] =]T(A®By). (S.l) Y/=i / y=i The proof is by mathematical induction. Result (S.l) is valid for s = 2, as is evident from result (1.12). Suppose now that result (S.l) is valid for s = s*. Then, making use of result (1.12), we find that ,(Eb^)=A0(Eb>+b^+i)
140 16. Kronecker Products and the Vec and Vech Operators = Uof^Byj + (A®B^+1) = £(A® By), /=1 which indicates that result (S.l) is valid for s = s* + 1, thereby completing the induction argument. Moreover, it can be shown in analogous fashion that, for any p x q matrix B, (Z!A')0B=Z!(A/ ® B) (S2) Now, making use of results (S.2) and (S.l), we find that (e a') ® (i>)=t [a' ® (i>)]=§x> ® »;>• EXERCISE 2. Show that, for any m x n matrix A and p x q matrix B, A ® B = (A ® I,,) diag(B, B B). Solution. Making use of results (1.20) and (1.7), we find that A ® B = (A ® 1,,)(1,, ® B) = (A ®lp) diag(B, B B). EXERCISE 3. Show that, for any m x 1 vector a and any p x 1 vector b, (1) a ® b = (a ® lp)b and (2) a' <g> b; = b'(a; ® Ip). Solution. Making use of results (1.20) and (1.1), we find (1) that a®b = (a®Ip)(l®b) = (a®Ip)b and similarly (2) that a' ® b; = (1 ® b;)(a; <g> lp) = b;(a; ® Ip). EXERCISE 4. Let A and B represent square matrices. (a) Show that if A and B are orthogonal, then A <g> B is orthogonal. (b) Show that if A and B are idempotent, then A <g> B is idempotent. Solution. Note that (since A and B are square) A ® B is square. (a) If A and B are orthogonal, then we have [in light of results (1.15), (1.19), and (1.8)] that (A ® B/(A ® B) = (A' ® B')(A ® B) = (A'A) ® (B'B) = I ® I = I
16. Kronecker Products and the Vec and Vech Operators 141 and hence that A ® B is orthogonal, (b) If A and B are idempotent, then we have [in light of result (1.19)] that (A®B)(A®B) = (AA)®(BB) = A®B and hence that A ® B is idempotent. EXERCISE 5. Letting //2,//, /?, and q represent arbitrary positive integers, show (a) that, for any p x q matrix B (having p > 1 or q > 1), there exists an m x n matrix A such that A ® B has generalized inverses that are not expressible in the form A~ ® B~ and (b) that, for any m x n matrix A (having m > 1 or n > 1), there exists apxq matrix B such that A ® B has generalized inverses that are not expressible in the form A~ ® B~. Solution, (a) Take A = 0. Then, A ® B = 0, so that any nq x mp matrix is a generalized inverse of A ® B. Since every one of the mn (q x p dimensional) blocks of the Kronecker product A~ ® B~ is a scalar multiple of the same q x p matrix (namely, B~), A®B has generalized inverses that are not expressible in the form A~ ® B~. Consider, for example, an nq x mp partitioned matrix comprising mn (q x p dimensional) blocks, including one block that has a single nonzero entry and a second block that also has a single nonzero entry but in a different location than the first. Clearly, this matrix is a generalized inverse of A ® B that is not expressible in the form A~ ® B~. (b) Take B = 0. Then, A ® B = 0, so that any nq x mp matrix is a generalized inverse of A ® B. Now, letting ct5 represent the tslh element of B~ and observing that (for t = 1,...,^ and s = 1 p) the n x m submatrix of A~ ® B~ obtained by striking out all of the rows and columns except the rth, (q + r)th [(« — 1)# 4- r]th rows and sth, (p + s)th [(//z - \)p + s]th columns equals ct5Ar, it follows that A ® B has generalized inverses that are not expressible in the form A~ ® B~. Consider, for example, an nq x mp matrix for which the n x m submatrix obtained by striking out (for some t and s) all of the rows and columns except the /th, (q 4- f)th [(/i — \)q + f]th rows and sth, {p + 5)th, ..., [(/// -1)//+s]th columns has a single nonzero entry and for which the n x m submatrix obtained by striking out (for some /' and s' with t' £ tors' ^ s) all of the rows and columns except the /;th, (q + r')th [(« - 1)# 4- /;]th rows and s'th, (p+s')th [(/// - l)//+5;]th columns also has a single nonzero entry but in a different location than the first submatrix. Clearly, this matrix is a generalized inverse of A ® B that is not expressible in the form A~ ® B~. EXERCISE 6. Let X = A ® B, where A is an m x n matrix and B a p x q matrix. Show that Px = PA ® PB. Solution. According to result (1.15), X; = A' ® B;. Thus, making use of result (1.19), we find that X'X = (A' ® B')(A ® B) = (A'A) ® (B'B),
142 16. Kronecker Products and the Vec and Vech Operators so that (A'A)~ ® (B'B)~ is a generalized inverse of X'X. And, again making use of result (1.19), it follows that Px = XiX'XyX' = (A ® B)[(A'A)~ ® (B'B)-](A' ® B') = [A(A'A)-A'] ® [B(B'B)-B'] = PA®Pb. EXERCISE 7. Show that the Kronecker product A ® B of an m x m matrix A and an n x n matrix B is (a) symmetric nonnegative definite if A and B are both symmetric nonnegative definite or both symmetric nonpositive definite and (b) symmetric positive definite if A and B are both symmetric positive definite or both symmetric negative definite. Solution, (a) Suppose that A and B are both symmetric nonnegative definite. Then, according to Corollary 14.3.8, there exist matrices P and Q such that A = VY and B = Q'Q. Thus, making use of results (1.19) and (1.15), we find that A®B = (P'fcQ'HPfcQ) = (P®Q)'(P®Q). We conclude (in light of Corollary 14.3.8 or 14.2.14) that A ® B is symmetric nonnegative definite. Alternatively, if A and B are both symmetric nonpositive definite, then -A and -B are symmetric and (by definition) nonnegative definite, and the proof [of Part (a)] is complete upon observing [in light of result (1.10)] that A ® B = (-A)®(-B). (b) Suppose now that A and B are both symmetric positive definite. Then, according to Corollary 14.3.13, there exist nonsingular matrices P and Q such that A = P/P and B = Q'Q. Further, A® B = (P® Q)'(P® Q), and P® Q is nonsingular. We conclude (in light of Corollary 14.3.13 or 14.2.14) that A ® B is symmetric positive definite. Alternatively, if A and B are both symmetric negative definite, then -A and -B are symmetric and (by definition) positive definite, and the proof is complete upon observing that A ® B = (-A) ® (-B). EXERCISE 8. Let A and B represent m x m symmetric matrices and C and D n x n symmetric matrices. Using the result of Exercise 7 (or otherwise), show that if A - B, C - D, B, and C are nonnegative definite, then (A ® C) - (B ® D) is symmetric nonnegative definite. Solution. Using properties (1.10) - (1.12), we find that (A®C)-(B®D) = {[(A-B)+B]®C}-{B®[C-(C-D)]} = [(A-B)®C] + (B®C) -{(B®C)-IB®(C-D)]} = [(A - B) ® C] + [B ® (C - D)]. (S.3)
16. Kronecker Products and the Vec and Vech Operators 143 Now, suppose that A - B, C - D, B, and C are nonnegative definite. Then, it follows from the result of Part (a) of Exercise 7 that (A - B) <g> C and B <g> (C - D) are both symmetric nonnegative definite and hence (in light of Lemma 14.2.4) that their sum [(A - B) <g> C] + [B <g> (C - D)] is symmetric nonnegative definite. And, based on equality (S.3), we conclude that (A <g> C) - (B <g> D) is symmetric nonnegative definite. EXERCISE 9. Let A represent an m x n matrix and B a p x q matrix. Show that, in the case of the usual norm, ||A®B||=||A|| ||B|| . Solution. Making use of results (1.15), (1.19), and (1.25), we find that IIA <g> B || = {tr[(A <g> B)'(A ® B)]}5 = {tr[(A'<g>B')(A®B)]}5 = {tr[(A'A)®(B'B)]}3 = {tr(A'A)}5 {tr(B'B)}2 = I|A||||B||. EXERCISE 10. Verify that, for an m x n partitioned matrix A = [An Ai2 A21 A22 Ari Ar2 and a p x q matrix B, A<g>B = /Au<g>B Aj2®B A2i®B A22®B Ajc ® B\ A2c®B \Ari ® B Ar2 ® B ... A,.c ® B/ that is, A ® B equals the mp x nq matrix obtained by replacing each block A,-y- of A with the Kronecker product of A/; and B. Solution. For 1 = 1 r, let m,- represent the number of rows in An, A/2 A,c; and, for j = 1 c, let nj represent the number of columns in Ajy, A2y\ ..., Ar;. Define F = A ® B; and partition F as /Fn F|2 ... F,c\ F21 F22 ... F2c F = \Frl Fr2 Frc/
144 16. Kronecker Products and the Vec and Vech Operators that is, partition F into /• rows and c columns of blocks, the //th of which is of dimensions nup x r\jq and is denoted by F;/. Then (for i = 1 r and 7 = 1 c), F,-; equals a partitioned matrix comprising /?z; rows and nj columns of p x q dimensional blocks, the z/uth of which is (When / = 1 or j = 1, interpret the degenerate sum wzj -\ \- m;-\ or /?i + h rij-\ as zero.) It follows that F// = A,y- <g> B. EXERCISE 11. Show that (a) if T and U are both upper triangular matrices, then T <g> U is an upper triangular matrix and (b) if T and L are both lower triangular matrices, then T ® L is a lower triangular matrix. Solution, (a) Suppose that T = {f,;} is an upper triangular matrix of order m and U = {iff/} an upper triangular matrix of order /z. Then, T ® U is a square matrix of order m/z, and the element that appears in the [n(i — 1) + r]th row and [/7(./ -1) + s]th column of T ® U is UjUrs- Clearly, Ujitrs ^ 0 only if j > i and s > r. Thus, the element that appears in the [n (/-1)+ r]th row and [n(j -1)+ s]th column of T <g> U is nonzero only if n (j - 1) + s > n (/ - 1) + r. It follows that T <g> U is an upper triangular matrix. (b) Suppose that T and L are both lower triangular matrices. Then, V and L' are both upper triangular, and consequently it follows from Part (a) that T' ® L' is upper triangular. Since [in light of result (1.15)] T ® L = (T; ® L')', we conclude that T ® L is lower triangular. EXERCISE 12. Let A represent an m x m matrix and B an n x n matrix. Suppose that A and B have LDU decompositions, say A = LiD|Ui and B = L2D2U2. Using the results of Exercise 11, show that A ® B has the LDU decomposition A®B = LDU,whereL = Lj ®L2,D = Dj ®D2,andU = Ui ®U2. Solution. That A ® B = LDU is an immediate consequence of result (1.19). Moreover, D is (by definition) the Kronecker product of two diagonal matrices (namely, D1 and D2) and hence is diagonal. And, U is (by definition) the Kronecker product of two upper triangular matrices (namely, Ui and U2) and hence [as a consequence of Part (a) of Exercise 11] is upper triangular. Similarly, L is (by definition) the Kronecker product of two lower triangular matrices (namely, L| and L2) and hence [as a consequence of Part (b) of Exercise 11] is lower triangular. It remains to show that the diagonal elements of L and U equal one. In this regard, observe that the [(« - 1)/ + /]th diagonal element of L is the product of the /th diagonal element of L| and the /th diagonal element of L2 and that the \(n -1)/ + r]th diagonal element of U is the product of the /th diagonal element of Ui and the rth diagonal element of U2. Since Li,L2,Ui, and U2 are unit triangular, their diagonal elements equal one. Thus, the diagonal elements of L and U equal one. EXERCISE 13. Let Ai, A2 A* represent k matrices (of the same dimen-
16. Kronecker Products and the Vec and Vech Operators 145 sions). Show that A|, Ai A* are linearly independent if and only if vec(A|), vec(A2) vec(Ajt) are linearly independent. Solution. It suffices to show that A|, A2, -.., A* are linearly dependent if and only if vec(Aj), vec(A2), • • • ♦ vec(Ajt) are linearly dependent. Suppose that Aj, A2 A* are linearly dependent. Then, there exist scalars fi, c2 Q. not all zero, such that £f=J cyA/ = 0. Since [in light of result (2.6)] * * y^cj vec(Aj) = vec(^c,A/) = vec(0) = 0, /=1 /=1 we conclude that vec(A|), vec(A2) vec(Ajt) are linearly dependent. Conversely, suppose that vec(A|), vec(A2) vec(Ajt) are linearly dependent. Then, there exist scalars c\, C2, -.., c*, not all zero, such that k y\/vec(Aj) =0, /=1 or equivalently [in light of result (2.6)] such that vec(£f=1 cyA/) = 0, and hence such that 5Zf=1 ct A/ = 0. We conclude that Aj, A2 A* are linearly dependent. EXERCISE 14. Let m represent a positive integer, let e,- represent the /th column of lm (/ = 1 m), and (for /, 7 = 1 m) let U,-; = e,e'y. (in which case 1¾ is an m x m matrix whose //th element is 1 and whose remaining m2 — 1 elements areO). (a) Show that m vec(Iw) = ^]e/®e/-. /=1 (b) Show that (for /, j, r, s = 1 m) vec(Ur,)[vec(U,;)]' = U/y ® U„ . (c) Show that 1» m EEU'V ® U'V = vec(I„,)[vec(I,H)]'. /=1 ;=i Solution, (a) Making use of results (2.4.4), (2.6), and (2.3), we find that vec(Im) = vec(^e/ej) = ^ vec(e,-e{) = ]Te/ ® e,-. / 1 / (b) Making use of results (2.3), (1.15), and (1.19), we find that vec(Ur/)[vec(U,i/)l' = vec(ere;.)[vec(e,e'y)]'
146 16. Kronecker Products and the Vec and Vech Operators = (ei<S>er)(ej®es)' = (e/<g>er)(e';<g><) = (eie'j) <g> ere's) = 1¾ <g> \Jrs. (c) Making use of Part (b) and results (2.6) and (2.4.4), we find that J^Uij ®UU = ^vec^O^tvec^)]' i,j i j = ^;vec(U/•l•)[^;vec(U77)], » j = vec(J]U/7)[vec(^;Uy7)], » j = vec(Im)[vec(Im)]'. EXERCISE 15. Let A represent an n x n matrix. (a) Show that if A is orthogonal, then (vec A)'vec A = n. (b) Show that if A is idempotent, then [vec(A/)]/vec A = rank(A). Solution, (a) If A is orthogonal, then, making use of result (2.14), we find that (vec A)'vec A = tr(A'A) = tr(I„) = /i. (b) If A is idempotent, then, making use of result (2.14) and Corollary 10.2.2, we find that [vec(A/)]/vec A = tr(AA) = tr(A) = rank(A). EXERCISE 16. Show that for any m x n matrix A,px« matrix X, p x p matrix B, and n x m matrix C, tr(AX'BXC) = (vec X)/[(A/C/) <g> B]vec X = (vec X)'[(CA) <g> B']vec X. Solution. Making use of results (5.2.3) and (2.15), we find that tr(AX'BXC) = tr(X'BXCA) = tr[X/BX(A/C/)/] = (vec X),[(A,C/) <g> B]vec X. Further, as a consequence of Lemma 14.1.1 and result (1.15), we have that (vec X),[(A,C) <g> B]vec X = (vec X),[(A,C) <g> Bfvec X = (vec X)/[(A/C/)/ ® B']vec X = (vecX)/[(CA)®B/]vecX. EXERCISE 17. (a) Let V represent a linear space of m x n matrices, and let g represent a function that assigns the value A * B to each pair of matrices A and B in V. Take U to be the linear space of mn x 1 vectors defined by U = {xe TZmnxl : x = vec(A) for some A e V},
16. Kronecker Products and the Vec and Vech Operators 147 and let x • y represent the value assigned to each pair of vectors x and y in U by an arbitrary inner product /. Show that g is an inner product (for V) if and only if there exists an / such that (for all A and B in V) A * B = vec(A)*vec(B). (b) Let g represent a function that assigns the value A * B to an arbitrary pair of matrices A and B in Kmxn. Show that g is an inner product (for 1Zmxn) if and only if there exists an mn x mn partitioned symmetric positive definite matrix W = w21 W12 W22 \Wnl W„; W2,, wmiy (where each submatrix is of dimensions m x m) such that (for all A and B in 1Zm xn) A*B = ^a;.Wiyb;, Ui where ai, a2 a„ and bj, D2 b„ represent the first, second nth columns of A and B, respectively. (c) Let g represent a function that assigns the value x' * y' to an arbitrary pair of (row) vectors in 1Zlxn. Show that g is an inner product (for 1Zlxn) if and only if there exists an n x n symmetric positive definite matrix W such that (for every pair of /i-dimensional row vectors x' and y') x^y^x'Wy. Solution, (a) Suppose that, for some /, A * B = vec(A)* vec(B) (for all A and B in V). Then, (1) A * B = vec(A) • vec(B) = vec(B) • vec(A) = B * A; (2) A * A = vec (A) • vec (A) > 0, with equality holding if and only if vec (A) = 0 or equivalently if and only if A = 0; (3) (k\) *B = vec(/:A)-vec(B) = [k vec(A)]-vec(B) = k [vec(A) • vec(B)] = k(\ * B); (4) (A + B) * C = vec(A + B) • vec(C) = [vec(A) + vec(B)]-vec(C) = [vec(A)-vec(C)] + [vec(B)-vec(C)] = (A*C) + (B*C)
148 16. Kronecker Products and the Vec and Vech Operators (where A, B, and C represent arbitrary matrices in V and k represents an arbitrary scalar). Thus, g is an inner product. ^ Conversely, suppose that g is an inner product, and consider the function / that assigns to each pair of vectors x and y in U the value x*y = X*Y, where X and Y are the (unique) m x n matrices such that x = vec(X) and y = vec(Y). Then, letting x, y, and z represent arbitrary vectors in U, taking X, Y, and Z to be m x n matrices such that x = vec(X), y = vec(Y), and z = vec(Z), and denoting by k an arbitrary scalar, we find that (1) x*y = X*Y = Y*X = y*x; (2) x*x = X*X>0, with equality holding if and only if X = 0 or equivalently if and only if x = 0; (3) (*x) *y = (JfcX) * Y = k(X * Y) = k(x*y); (4) (x + y)*z = (X + Y)*Z = (X*Z) + (Y*Z) = (x*z) + (y*z). Thus, f is an inner product (for U). Moreover, for / = /, we have that A * B = vec(A) * vec(B) = vec(A)*vec(B) (for all A and B in V). (b) Let / represent an arbitrary inner product for TZm"x l, and let x*y represent the value assigned by / to an arbitrary pair of mn-dimensional column vectors x and y. According to Part (a), g is an inner product (for 71"'x") if and only if there exists an / such that (for all A and B in V) A * B = vec(A)*vec(B). Moreover, according to the discussion of Section 14.10a, every inner product for 7£""'xl is expressible as a bilinear form, and a bilinear form (in mn x 1 vectors) qualifies as an inner product for ft"'"xl if and only if the matrix of the bilinear form is symmetric and positive definite. Thus, g is an inner product (for 1Z'"X") if and only if there exists an mn x mn symmetric positive definite matrix W such that (for all A and B in V) A*B = (vecA);WvecB. Further, partitioning W as W = W2i w22 Vw„, w„2 ... W|„\ ... Wo,/ w„„/
16. Kronecker Products and the Vec and Vech Operators 149 and denoting by aj, a2 a„ and bj, D2 b„ the first, second /zth columns of A and B, respectively, we find that (vec A)' W vec B = ]TajW/yb; . 'J (c) It follows from Part (b) that g is an inner product (for TZlx") if and only if there exists an n x n symmetric positive definite matrix W = {u»/y} such that (for every pair of//-dimensional row vectors x7 = {.v,} and y7 = {y,}) x,*y, = ^2xiwijyj i-j or equivalently such that (for every x7 and y7) x7*y7 = x7Wy. EXERCISE 18. (a) Define (for m > 2) P to be the mn x mn permutation matrix such that, for every m x n matrix A, (vec A*\ _ . r *j=PvecA, where A* is the (wi — 1) x n matrix whose rows are respectively the first, ..., (/// — 1 )th rows of A and r7 is the /77th row of A [and hence where A = I J* J and A7 = (A;,r)]. (1) ShowthatK,,,M = (Ky'H JJp. (2) Show that |P| = (_i)0»-i>»<»-i>/2. (3) Show that |K,„„| = (-l)('"-1,"(n-n/2|K,„_I,„|. (b) Show that |K,„„| = (_i)«c»-i)ii(»-i)/4 (c) Show that |K,WH| = C_i)'»(»—i)/2. Solution, (a) (1) Since vec (AJ,) = K,„_i.„ vec A*, we have [in light of the defining relation (3.1)] that /%„_,.„ OWvecAA (V t,)"-*- K,„„vec A = vec (A7) = I r ) = Thus, K,,„,a=(KV'"
150 16. Kronecker Products and the Vec and Vech Operators for every mn-dimensional column vector a, implying (in light of Lemma 2.3.2) that V _ (Km-l.i? 0 \ p (2) The vector r is the n x 1 vector whose first, second, ..., «th elements are respectively the /nth, (2m)th (nm)th elements of vec A, and vec A* is the (mi — l)/i-dimensional subvector of vec A obtained by striking out those n elements. Accordingly, P = (' I, where P2 is the n x mn matrix whose first, second,..., wth rows are respectively the wth, (2m)th,..., («m)th rows of Imn, and Pi is the (m — \)n x mn submatrix of Imn obtained by striking out those n rows. Now, applying Lemma 13.1.3, we find that |P| = (-1)^, where 0 = (m - 1)1 + (m - 1)2 + • • • + (m - 1)(/2 - 1) = (m - l)n(n - 1)/2. (3) It follows from Parts (1) and (2) that |K„J = |P| 0 l„ = (-l)im-mn-l)/2\Km-l.nl (b) It follows from Part (a) that, for i > 2, IKf.1 = (-1)^^-^1^-1^1. By applying this equality m — 1 times (with i = m, m — 1,..., 2, respectively), we find that, for m > 2, |K,„„| = (-1)^-^-^11^..,.,,1 = (_1)C—ll«C»-l)/2(_1)C»-2)«C«-ll/2|Kiii__2_B| = (_l)[<'"-l)+(ro-2)+"H-l]M(ii-l)/2|K I _ (.^^^-1)/211.(1.-1)/2^^1 = (_l)'»('»-l)«(i»-l)/4|Kirt| Since Kj,, = ln [and since (-1)° = 1], we conclude that (for //? > 1) \Knu,\ = (-\)m{m-l)"{n-lV*. (c) It follows from Part (b) that 1^,,,1 = (-1)^-1^2. Since the product of two odd numbers is odd and the product of two even numbers even, we conclude that 1^,,,,,1 = (-1)^-^/2. EXERCISE 19. Show that, for any /// x /1 matrix A, p x 1 vector a, and q x 1 vector b,
16. Kronecker Products and the Vec and Vech Operators 151 (1) b' <g> A <g> a = Kmp[(ab') <g> A]; (2) a <g> A ® b' = Kpm [A <g> (ab')]. Solution. Making use of Corollary 16.3.3 and of results (1.16) and (1.4), we find that b' <g> A <g> a = (b' <g> A) <g> a = K„p [a <g> (b' <g> A) ] = Kmp[(a <g> b') <g> A] = Kmp[(ab') <g> A] and similarly that a<g>A<g>b' = a<g>(A<g>b') = Kpm[(A®b') <g>a] = Kpm[A <g> (b' <g> a)] = Kpm[A <g> (ab')]. EXERCISE 20. Let /w and n represent positive integers, and let e,- represent the i th column of Im (i = 1 m) and uy represent the jth column of 1„ (j = 1 /i). Show that n m Km„ = J^uy ® Im ® uy = J^e,- ®I„ ® ej. y=i «=i Solution. Starting with result (3.3) and using results (1.4) and (2.4.4), we find that = ^Uy®e,-®eJ®uy «./ y » = £]u'y ® (^e/ej) ®uy = £]u'; ® Im ® Uy . Similarly, Kwn = ]£ (fiiu'j) ® (UyeJ) i-j = 5^el-®u/y®uy®e; = 5^ e,- ® (]Puy ® uy) ® ej » y = IZe»- ® <ZXu'y> ® e'i = l]e/ ® In ® ej. EXERCISE 21. Let /w,«, and p represent positive integers. Using the result of Exercise 20, show that
152 16. Kronecker Products and the Vec and Vech Operators (a) IV,„piW = &■ p,VItl**■>!!,Up \ (D) iVmp.»iv»p,m^////.p = 11 (C) t^tl,mp = **H/MH*»7HM.p I (d) IVp,;»;jtVw<;jp = *V»Mip<&p,m» "» (e) JV„p.,},lV,„„.p = ^-IHH.p^-Hp.lH "» (I) **m.lip***Hp.M = «7H/>.M **»»,»/> • [//i/tf. Begin by letting u, represent the jth column of I„ and showing that KmPm„ = 5Zy (u} ® Ip) ® (In, ® u;) and then making use of the result that, for any m x n matrix A and p x q matrix B, B <g> A = KP,„(A ® B)K„g .] Solution, (a) Letting uy represent the jth column of I„ and making use of the result of Exercise 20 and the result cited in the hint [or equivalently result (3.10)], we find [in light of results (1.8), (1.16), (1.4), and (2.4.4)] that K,„p.n =^Uy<g>I,„p<g>U; J j = ^(^.01,,)0(1,,, ®Uj) j = ^Kp,,„„[(I„, <g> Uj) <g> (Uy <g> lp)]Kmj,p j = ^K/M„„[I„, ® (u;uy> ® JpIKm.Hp j = 1^.,,,,,(1,,, <g> (^uyu}) <g> Ip]K„,,„p i = Kp>H,M (1,,, ® I„ ® Ip)K,i,,„p 1=1 &p.mn*mnp'*m,np = *vp,nn»**iw.»p • (b) Making use of Part (a) and result (3.6), we find that **mp.n **#»/>,«» '^■mn.p = *>/>.»i»»»*Vm.Hp*v»p.m*VmM,p = *^p,mn I K»hi,/? = Kptii,„K„„,tP = I. (c) Making use of Part (b) and result (3.6), we find that &-n,nip = **n.mp*mnp = **-n.mp"*np,n'*np.m**mn,p = 1 &itp.m**mn.p = **•»/>.!»**•»»/»./> ■ (d) Using Part (a) (twice), we find that &p.milKm.lip = Kf»;>,» = K,,j„.„ = Kf|,.y)|,Kp>„„; = K,„.,,pKp>„„, .
16. Kronecker Products and the Vec and Vech Operators 153 (e) Using Part (c) (twice), we find that **»p,/M*VMii,p = &n.mp = Kfi.pm = K-nm.p'&pn.m = K»i/i.pKMp,in • (f) Making use of Parts (c) and (a) and result (3.6), we find that **7l!,»P**7Hp.H = (Knip<nK}l||l|p)(Kp.}|lnKnl>np) EXERCISE 22. Let A represent an m x n matrix, and define B = Kmn(A' ® A). Show (a) that B is symmetric, (b) that rank(B) = [rank(A)]2, (c) that B2 = (AA') <g> (A'A), and (d) that tr(B) = tr(A'A). Solution, (a) Making use of results (1.15), (3.6), and (3.9), we find that B' = (A' ® A)'K;„, = (A ® A')K,„„ = Km„(A' <g> A) = B. (b) Since Kmn is nonsingular, rank(B) = rank(A' ® A). Moreover, it follows from result (1.26) that rank(A' <g> A) = rank(A') rank(A). Since rank(A') = rank(A), we conclude that rank(B) = [rank(A)]2. (c) Making use of results (3.10) and (1.19), we find that B2 = Kmw(A' <g> A)KW„(A' <g> A) = (A <g> A')(A' <g> A) = (AA') <g> (A'A). (d) That tr(B) = tr(A'A) is an immediate consequence of the second equality in result (3.15). EXERCISE 23. Show that, for any m x n matrix A and any pxq matrix B, vec(A <g> B) = (I„ <g> G)vec A = (H <g> Ip)vec B, where G = (K^,,, ® Ip)(I,„ ® vec B) and H = (Iw ® K9,„)[vec(A) ® 1^]. Solution. Making use of results (1.20), (1.1), and (1.8), we find that vec(A) <g> vec(B) = (Iw„ <g> vec B)[vec(A) <g> 1] = (I/n» ® vec B)vec A = (I„ <g> I,„ ® vec B)vec A and similarly that vec(A) <g> vec(B) = [vec(A) <g> Ip9](l <g> vec B) = [vec(A) <g> Ip^]vec B = [vec(A) ®lq® Ip]vec B. Now, substituting these expressions [for vec(A) <g> vec(B)] into formula (3.16) and making use of result (1.19), we obtain vec(A <g> B) = [I„ ® (Kqm <g> Ip)][I„ <g> (lm <g> vec B)]vec A = (IM <g> G)vec A
154 16. Kronecker Products and the Vec and Vech Operators and vec(A <g> B) = [(1,, <g> Kg„,) <g> Ip]{[vec(A) <g> lq] <g> Ip}vec B = (H <g> Ip)vec B. EXERCISE 24. Show that, for H„ = (G^Cr'G;,, G,tH„H„ = H„ . Solution. ForH, = (G;,G„)-lG;,, G„ H„ H„ = G„ (GM G„)" GM G„ (G„ G„) = G„ (G„ G„) = H„ . EXERCISE 25. There exists a unique matrix L„ such that vech A = L„ vec A for every n xn matrix A (symmetric or not). [The matrix L„ is one choice for the matrix H,„ i.e., for a left inverse of G„. It is referred to by Magnus and Neudecker (1980) as the elimination matrix—the effect of premultiplying the vec of an n x n matrix A by L„ is to eliminate (from vec A) the "supradiagonal" elements of A.] (a) Write out the elements of Lj, L2, and L3. (b) For an arbitrary positive integer /2, describe L„ in terms of its rows. Solution, (a) Lj = (1), /1 0 0 0\ L2= 0 1 0 0 , \0 0 0 1/ and L3 = /1 0000000 0\ 010000000 001000000 000010000 000001000 \o 0000000 (b) For i > j, the [(j - 1)(// - j/2) + /]th row of L„ is the [(j -1)//+ /]th row of In2. EXERCISE 26. Let A represent an // x // matrix and b an // x 1 vector. (a) Show that (1/2)[(A <g> b;) + (b; <g> A)]G„ = (A <g> b;)G„. (b) Show that, for H„ = (G^G,,)"^;,, (1)(1 /2)H„[(b ® A) + (A ® b)l = H„ (b ® A); (2) (A <g> b')G„H„ = (1/2)[(A ® b') + (b' <g> A)]; (3) G„H„(b®A) = (l/2)[(b®A) + (A®b)l.
16. Kronecker Products and the Vec and Vech Operators 155 Solution, (a) Using results (3.13) and (4.16), we find that (1/2)[(A <g> b') + (b' <g> A)]G„ = (A ® b')[(l/2)(Iw2 + Kn„)]G„ = (A <g> b')G„. (b) Using results (3.12), (4.17), (4.22), and (3.13), we find that, for H„ = (G^cr'G;, (1) (l/2)H„[(b<g>A) + (A<g>b)] = H„[(l/2)(IM2+Km,)](b<g>A) = H„(b<g>A); (2) (A®b')G(lH„ = (A®b')[(l/2)(In2 + K„„)] = (1/2)[(A®b') + (b'<g>A)]; (3) GnH„(b <g>A) = (1/2)(1,,2 + KHH)(b <g> A) = (l/2)[(b<g> A) + (A <g>b)]. EXERCISE 27. Let A = {a,j} represent annxn (possibly nonsymmetric) matrix. (a) Show that, for H„ = (G^GJ-'G;,, H„ vec A = (1/2) vech(A + A'). (b) Show that G^G„ vech A = vech[2A - diagfan, an ann)]. (c) Show that G^ vec A = vech[A + A' - diag(a i i, 022 a„n)]. Solution, (a) Since (1/2) (A + A') is annxn symmetric matrix, we have [in light of result (4.17)] that, for H„ = {G'nG„rlG'„, H„ vec A = H„[(l/2)(In2 + K„„)]vec A = (l/2)Hn(vecA + KimvecA) = (l/2)H„(vecA + vecA') = (l/2)H„vec(A + A') = (l/2)HnG„ vech(A + A') = (l/2)vech (A + A'). (b) The matrix GJ,G„ is diagonal. Further, the [(/ - l)(n - //2) + /]th diagonal element of GJ,G„ equals 1, and the [(/ - 1 )(n - //2) + /]th elements of vech A and vech[2A-diag (an, a22 flw,)]bothequalal/,sothatthe[(/-l)(n-//2)+/]th elements of GJ,G„ vech A and vech [2A - diag («11, «22 am)\ both equal an. And, for i > ;, the [(; - l)(n - j/2) + /]th diagonal element of G;,Gn equals 2, the [{j - 1)(« - j/2) + /]th element of vech A equals a/y, and the [(; — l)(n — j/2) + i]th element of vech [2A - diag («11, «22 a„„)] equals 2a,j, so that (for i > ;) the [(j - \)(n - j/2) + /]th elements of G^G„ vech A and vech[2A - diag (a\ 1, «22 a„„)] both equal 2ayr We conclude that GJ,G„ vech A = vech[2A - diag (a\ \, a22 a„„)].
156 16. Kronecker Products and the Vec and Vech Operators (c) Using the results of Parts (a) and (b), we find that G; vec A = GiCkKGiCr'G; vec A] = (l/2)G;,G„vech(A + A') = (l/2)vech[2(A + A') - diag (2fljj, 2a22 2a„„)] = vech[A + A' — diag (an, «22 o„„)]. EXERCISE 28. Let A represent a square matrix of order n. Show that, for H„ = g„h„(A®a)h;, = (a®a)h;. Solution. Using result (4.26), we find that, for H„ = (G'nG„)~lG',r G„H„(A <g> A)H; = G„H„(A <g> A)G„(G,„G„rl = (A <g> AjG^G,,)-1 = (A®A)H;,. EXERCISE 29. Show that if an n x n matrix A = {a,;} is upper triangular, lower triangular, or diagonal, then H„(A <g> A)G„ is respectively upper triangular, lower triangular, or diagonal with diagonal elements an ays (/ = 1 //; j = / //). Solution. Let us use mathematical induction to show that for any // x // upper triangular matrix A = {a,;}, H„(A ® A)G„ is upper triangular with diagonal elements aucijj (/ = 1 //; j = i //). For every 1 x 1 upper triangular matrix A = (a\\), Hj(A®A)Gi is the 1 x 1 matrix (a2,), which is upper triangular with diagonal element auajj (/ = 1; j = 1). Suppose now that, for every // x//upper triangular matrix A = {oij}, H„(A®A)G„ is upper triangular with diagonal elements fl,/fly; (/ = 1 //; j = / //),and let B = [bjj} represent an (n +1) x (11 +1) upper triangular matrix. Then, to complete the induction argument, it suffices to show that H„+i (B <g> B)G„+i is upper triangular with diagonal elements bnbjj (/ = 1,..., n + 1; j = /,..., n + 1). For this purpose, partition B as -c a (where A is n x n with //th element bj+ij+\). Then (since B is upper triangular), a = 0, and it follows from result (4.29) that (c- 2cb' (b'®b')GM \ 0 rA (b'®A)G„ . 0 0 H„(A<g>A)G„/ Moreover, A is upper triangular, and hence (by supposition) H„ (A® A)G„ is upper triangular with diagonal elements bubjjii = 2 // + U j = / // + 1).
16. Kronecker Products and the Vec and Vech Operators 157 Thus, H„+i (B ® B)G„+| is upper triangular. And, its diagonal elements are c2 = bu' cbJj = b\\bjj(j =2 «+1), andbiibjjd = 2 n+\\j = i «+ 1); that is, its diagonal elements are bubjjii = 1 n + 1; j = i n + 1). It can be established via an analogous argument that, for any n x n lower triangular matrix A = {a,y},H„(A <g> A)G„ is lower triangular with diagonal elementsaaajj (/ = 1 «; j =i n). Finally, note that if an n x n matrix A = {atj} is diagonal, then A is both upper and lower triangular, in which case H„(A <g> A)GM is both upper and lower triangular, and hence diagonal, with diagonal elements aaajj (/ = 1,...,»; j = / n). EXERCISE 30. Let Aj A*, and B represent m x n matrices, and let b = vecB. (a) Show that the matrix equation Yll=\ A*'^» = B (in unknowns .yj xt) is equivalent to a linear system of the form Ax = b, where x = (x\ a*)' is a vector of unknowns. (b) Show that if Aj A*, and B are symmetric, then the matrix equation 5Zf=i *iAi = B (in unknowns jci, ..., a*) is equivalent to a linear system of the form A*x = b*, where b* = vechB and x = (aj, ..., jr*)' is a vector of unknowns. Solution. Let A = (vec Aj vec A*). (a) Making use of result (2.6), we find that vec(Y^ a,-A,) = y^v,- vec A,- = Ax. i i Since clearly the (matrix) equation J^i xi^i = B is equivalent to the (vector) equation vec(£,- a/A,) = vec B, we conclude that the equation £,- .v,A/ = B is equivalent to the linear system Ax = b. (b) Suppose that Aj A*, and B are symmetric (in which case m = n). And, let A* = (vech Aj,..., vech At). Then, for any value of x such that A*x = b*, Ax = G„A*x = G„b* = b, and conversely, for any value of x such that Ax = b, A*x = H„Ax = H„b = b*. We conclude that the linear system A*x = b* is equivalent to the linear system Ax = b and hence [in light of Part (a)] equivalent to the equation £,- a/A,- = B. EXERCISE 31. Let F represent a p x p matrix of functions, defined on a set 5, of a vector x = (aj a,,,)' of m variables. Show that, for k = 2,3
158 16. Kronecker Products and the Vec and Vech Operators (where F° = lp). Solution. Making use of results (6.1), (15.4.8), and (2.10), we find that d vec(F*) _ f3(F*)1 dXj " vec[ dxj J = ^ Ffc-> |* + F*-2|^F + • ■ • + ^Fk~l) \ dxj dxj dxj ) = vec(gF^-) implying that •V=l EXERCISE 32. Let F = [fiS] and G represent p x q and r x s matrices of functions, defined on a set 5, of a vector x = (x\,..., x,„)' of m variables. (a) Show that (for./ = 1 m) 3(F®G) F®G) / 3G\ /3F A = 1 m) 7®G) „ „ . X, ™ 3vecG 3vecF „1 = (Iq <g> Ksp <g> Ir) (vecF) <g> — + — <g> (vecG) . 7 L °xj OX) J (b) Show that (for ; = 1 m) 8 vec(F ® G) 3.V; (c) Show that 3vec(F®G) „ „ w T ^ 3vecG 3vecF „1 — = dq ® Ksp <g> Ir) I (vec F) <g> -^- + -^- <g> (vec G) . (d) Show that, in the special case where x' = [(vec X)', (vec Y)'], F(x) = X, and G(x) = Y for some p x q and r x s matrices X and Y of variables, the formula in Part (c) simplifies to 3 vec(X <g> Y) — = (If/ <g> Ksp <g> I,.)[Ir<7 ® (vec Y), (vec X) <g> Ir.v].
16. Kronecker Products and the Vec and Vech Operators 159 Solution, (a) Partition each of the three matrices 3(F<g> G)/dxj, F<g> (dG/dxj)y and (dF/dxj) ® G into p rows and q columns ofrxs dimensional blocks. Then, for i = 1 p and s = 1 ^, the /5th blocks of 3(F <g> G)/dxj, F ® (dG/dxj), and (dF/dxj) ®G are respectively d{fisG)/dxh fis(dG/dxj), and (Bfis/Bxj)Gt implying [in light of result (15.4.9)] that the isth block of 3(F ® G)/3jcy- equals the sum of the isth blocks of F <g> (dG/dxj) and (dF/dxj) ® G. We conclude that 3(F®G) F®G) / 3G\ /3F A (b) Making use of Part (a) and Theorem 16.3.5, we find that 3 vec (F <g> G) dx: ra(F®G)"| =vec(F0S+vec(^0G) = (lq®Ksp®lr) (vec F)®vec(S)+vec(S)®(vecG)] o,]. „v 3vecG 3vecF , (vec F) ® — 1 —— ® (vec dxj dxj (c) In light of result (1.28), (vec F) ® [3(vec G)/dxj] is the ;th column of (vec F) ® [3(vec G)/dx!\. And, in light of result (1.27), [3(vec F)/3jcy] ® (vec G) is the ;th column of [3(vec F)/3x'] ® (vec G). Thus, it follows from Part (b) that 3 vec dx' (d) In this special case, 3vecG (F ® G) T B vec — = (lq ® Ksp ® lr) (vecF) ® —^ G d vec F , + __0(vec 4 dx' = (0, ir,), 8vecF dx' = (W 0), implying [in light of results (1.28) and (1.27)] that (vec F) ® -^- = [0, (vecF)<g>Ir5] dx' and that 3vecF ) (vec G) = [I„g ® (vec G), 0], so that it follows from Part (c) that 3 vec (F ® G) dx! = (lq® Ksp ® Ir)[lpq ® (vec G), (vec F) ® IrJ].
17 Intersections and Sums of Subspaces EXERCISE 1. Let U and V represent subspaces of 11"'x". (a) Show that WUVcW + V. (b) Show that U + V is the smallest subspace (of Hmx") that contains WUV, or, equivalently [in light of Part (a)], show that, for any subspace W such that UUV C W, W + V C W. Solution, (a) Let A represent an arbitrary /// x n matrix in U U V, so that (by definition) A e U or A e V. Upon observing that A = A + 0 = 0+Aand that the m x « null matrix 0 is a member of V and also oft/, we conclude that A e U+V. (b) Let A represent an arbitrary matrix in U + V. Then, A = U + V for some matrix V eU and some matrix V e V. Moreover, both U and V are in U U V, and hence both are in W (where W is an arbitrary subspace such that WUVC W). Since W is a linear space, it follows that A (= U + V) is in W. EXERCISE 2. Let A = 0 1 and B = 1 1 . Find (a) a basis for C(A) + /1 0\ /0 2\ ,= 0 1 andB= 1 1 l.R: \0 O) \2 l) C(B), (b) a basis for C(A) fl C(B), and (c) a vector in C(A) + C(B) that is not in C(A) U C(B). Solution, (a) According to result (1.4), C(A) + C(B) = C(A, B). Since the partitioned matrix (A, B) has only 3 rows, its rank cannot exceed 3. Further, the first 3 columns of (A, B) are linearly independent. We conclude that rank(A, B) = 3 and that the first 3 columns of (A, B), namely, (1,0,0)', (0,1,0)', and (0,1,2)',
162 17. Intersections and Sums of Subspaces form a basis for C(A, B) and hence for C(A) + C(B). (b) The column space C(A) of A comprises vectors of the form (*i, *2,0/ (where x\ and X2 are arbitrary scalars), and C(B) comprises vectors of the form (2.V2, yi +yi> 2>i +3>2)' (where y\ and )¾ are arbitrary scalars). Thus, C(A)flC(B) comprises those vectors that are expressible as (2)¾. yi +)>2> 2y\ +3)¾)7 for some scalars y\ and )¾ such that 2yj + 3y2 = 0 or equivalently (since 2yj + 3)¾ = 0 «#> )¾ = —2y\/3) of those vectors that are expressible as (-4yi/3, yi/3,0)' [= vi (-4/3,1/3,0)'] for some scalar y{. We conclude that C(A) fl C(B) is of dimension one and that the set whose only member is (—4, 1,0)' (obtained by setting vi = 3) is a basis for C(A) n C(B). (c) In light of the solution to Part (a), it suffices to find any 3-dimensional column vector that is not contained in C(A) or C(B). It follows from the solution to Part (b) that the vector (2)¾. c, 2vi + 3v2)\ where vi, V2, and z are any scalars such that 2yi + 3y2 # 0 and z i=- y\ +yi, is not contained in C(A) or C(B). For example, the vector (0,2,2)' (obtained by taking y\ = 1, )¾ = 0, and z = 2) is not contained inC(A)orC(B). EXERCISE 3. Let U, W, and X represent subspaces of a linear space V of matrices, and let Y represent an arbitrary matrix in V. (a) Show (1) that if Y X W and Y X Xy then Y X (W + X\ and (2) that if WlWand ULX.ihzn U±(W + X). (b) Show (I) thsx (U + W)±=U±HW± 2^6 {2) ±at(UnW±=U± + WL. Solution, (a) (1) Suppose that Y X W and Y ± X. Let Z represent an arbitrary matrix in W + X. Then, there exists a matrix W in W and a matrix X in X such that Z = W + X. Moreover, Y X W and Y X X, implying that Y X Z. We conclude that Y X (W + X). (2) Suppose that U X W and U X X. Let U represent an arbitrary matrix in U. Then, U X W and U X Xy implying [in light of Part (1)] that U X (W + X). We conclude that U ±(W + X). (b) (1) Observing that U C (U + W) and W C {U + W) and making use of Part (a)-(l), we find that YelW + W)1 <s> YKW + W) <S> YlWandYlW & Ye^andYeW1 <S> Yei^nW1). We conclude that (U + W)-1- =U±n W1. (2) Making use of Part (1) and Theorem 12.5.4, we find that u1 + wx = [cw1 + w1)1]1 = au1)1 n (W1)1!1 = (// n i n1.
17. Intersections and Sums of Subspaces 163 EXERCISE 4. Let Uy W, and X represent subspaces of a linear space V of matrices. (a)Show that{U nW) + (U n X) C Un(W + X). (b) Show (via an example) that U fl W = {0} and U C\ X = {0} does not necessarily imply that U fl (W + X) = {0}. (c) Show that ifWcW, then (1) U + W = U and (2) U fl (W + #) = W + (Wn#). Solution, (a) Let Y represent an arbitrary matrix in (U n W) + (ZY fl #). Then, Y = W + X for some matrix W in U fl W and some matrix X in U fl X. Since both W and X are in 14, Y is in Uy and since W is in W and X in X, Y is in W + X. Thus, Y is in U fl (W + #). We conclude that (UnW) + (UnX) C Un(W + X). (b) Suppose that V = TZlx2 and that Uy W, and X are the one-dimensional subspaces spanned by (1,1), (1,0), and (0,1), respectively. Then, clearly, U fl W = {0} and U fl X = {0}. However, W + # = ftlx2, and consequently wn(W+^)=w#{0}. (c) Suppose that W C W. (1) Since clearly U CU + W,it suffices to show that U + W C W. Let Y represent an arbitrary matrix in U + W. Then, Y = U+W for some matrix U in U and some matrix W in W. Moreover, W is in U (since WcW), and consequently Y is in U. We conclude that U + W C W. (2) It follows from Part (a) (and the supposition that WcW) that W+(UC\X) C Wn(W + #). Thus, it suffices to show that U fl (W + X) c W + (W n #). Let Y represent an arbitrary matrix in U fl (W + #). Then, Y eW + X,so that Y = W + X for some matrix W in W and some matrix X in Xy and also Y e U. Thus, X = Y - W, and, since (in light of the supposition that W C U) W (like Y) is in U, X is in U (as well as in X) and hence is in U n X. It follows that Y is in VV+ (^0^. We coiiclude that Wn(W + #) C W + (UnX). EXERCISE 5. Let U\,U2 Uk represent subspaces of TZmxn. Show that if, for j = 1,2 k, Uj is spanned by a (finite nonempty) set of (;w x #i) matrices U(/} U<f,then wi+%+■••+½=sP(u(1,) u^u;2) ie uf> u<*>>. Solution. Suppose that, for j = 1,2 /:, Wy is spanned by the set {Ujy) itff}. The proof that U\ +U2 + • • • + Uk is spanned by U^ U™, U(,2) Ujf Uf} Ujf is by mathematical induction. It follows from Lemma 17.1.1 that W1+W2 = sp(U(1,) U<;\U<2) 1¾¾).
164 17. Intersections and Sums of Subspaces Now, suppose that (for an arbitrary integer j between 2 and /:-1, inclusive) Wl+W2 + ...+W; = Sp(U;» umuci u(2) vu) rfrj))t Then, the proof is complete upon observing (in light of Lemma 17.1.1) that U\ +^2 + --- i-Uj+i = (Ui+U2 + ---+Uj)+Uj+i = spOJ^ U<[>, U<2> U<? U(/+1> tf£»). EXERCISE 6. Let U\ Uk represent subspaces of llmx". The k subspaces U\ Uk are said to be independent if, for matrices Uj 6 t/i,..., Ujt 6 £4, the only solution to the matrix equation Ui+-.- + 1¾ =0 (E.1) isU|=.'- = U*=0. (a) Show that U\ Uk are independent if and only if, for / = 2 k. Ui and Wj H h W/_i are essentially disjoint. (b) Show that U\ Uk are independent if and only if, for / = 1 k, Ui and Wj H h W,-i+ W,+i -\ \-Uk are essentially disjoint. (c) Use the results of Exercise 3 [along with Part (a) or (b)] to show that if U\,..., Uk are (pairwise) orthogonal, then they are independent. (d) Assuming that U\Mi Uk are of dimension one or more and letting {U^' Uj/*} represent any linearly independent set of matrices in Uj {j — 1,2 k), show that ifUiMi Uk are independent, then the combined set [V\]) U^, Uj2) U^2) V\k) Ujf) is linearly independent. (e) Assuming that U\,Ui Uk are of dimension one or more, show that U\, Uo Uk are independent if and only if, for every nonnull matrix Uj in U\, every nonnull matrix U2 in Ui and every nonnull matrix U* in Uk, Uj, U2 Ujt are linearly independent. (0 For j = 1 k, let pj = dim(t//), and let Sj represent a basis for Uj (j = 1 k). Define S to be the set of 5Z/=i Pj matrices obtained by combining all of the matrices in S\ 5* into a single set. Use the result of Exercise 5 [along with Part (d)] to show that (1) if U\ Uk are independent, then 5 is a basis for U\ +... +Uk\ and (2) if U\ Uk are not independent, then S contains a proper subset that is a basis for U\ +... + Uk- (g) Show that (1) if U\ Uk are independent, then dim(Wj + • • • + £4) = dimtfVi ) + ••• + dim{Uk); and (2) if U\ Uk are not independent, then dim(Wj +•••+#*)< dim(Wj ) + ••• + dim(Mjt).
17. Intersections and Sums of Subspaces 165 Solution, (a) It suffices to show that U\ Uk are not independent if and only if, for some i (2 < / < &), U\ and U\ -\ h Ut-\ are not essentially disjoint. Suppose that U\,..., Uk are not independent. Then, by definition, equation (E. 1) has a solution, say Ui = U*,..., U* = U£, other than Uj = • • • = U* = 0. Let r represent the largest value of / for which U* is nonnull. (Clearly, /* > 2.) Since uj + ---+u; = ui + ---+uj = o, U,* = -U* + • • • + (-Ur*_,) eUx + • • • + Ur-i. Thus, for i = r, Ui and U\ -\ \- Uj-\ are not essentially disjoint. Conversely, suppose that for some /\ say i = s, Ui and U\ -\ VUi-\ are not essentially disjoint. Then, there exists a nonnull matrix Vs such that Vs e Us and Vs eU\-\ \rUs-\. Further, there exist matrices Ui e U\ \}s-\ e Us-\ such that U.v = Ui + • • •+U5_i or equivalently such that Ui + • • •+\ls-\ + (~VS) = 0. Thus, equation (E.1) has a solution other than Ui = • • • = U* = 0. (b) It suffices to show that U\t...Mk are not independent if and only if, for some /(1 < i < &), Ui and U\ -\ YUi-\ +Ut+\ -\ 1-¼ are not essentially disjoint. Suppose that U\,...Mk are not independent. Then, by definition, equation (E.1) has a solution, say Ui = UJ U* = UJ, other than Uj = • • • = Ujt = 0. Let r represent an integer (between 1 and &, inclusive) such that U* ^ 0. Since U* = - £l;4rU*,U* is in the subspace 52/^ W/. as well as the subspace Ur. Thus, for i = /\ Ui and U\-\ V Ui-\ + Ui+\ -\ h Uk are not essentially disjoint. Conversely, suppose that for some /, say / = s, Uj and U\ -\ \-Ui-i +W,+i + —\-Uk are not essentially disjoint. Then, there exists a nonnull matrix U* such that Vs e Us and Us e J^i^s Ui. Further, there exist matrices Uj e U\ U5_i e U5-\, Ws+\ 6 Us+i Ujt e Uk such that U5 = ^,-^ U,- or equivalently such thatUi + • • • + Us-i + (-U*) + Vs+i + - • • + U* = 0. Thus, equation (E.1) has a solution other than Uj = ■ • = Ujt = 0. (c) Suppose that U\ Uk are orthogonal. Then, applying the result of Part (a)-(2) of Exercise 3(/-2 times), we find that U\ and U\ -\ h Ut-\ are orthogonal, implying (in light of Lemma 17.1.9) that Ut and U\-\ V Ui-\ are essentially disjoint (/=2 k). Based on Part (a), we conclude that U\ Uk are independent. (d) Suppose that U\, U2,..., Uk are independent. The proof that the set [\]\l\ .... Vlr\\ U(j2\ ..., U£\ ..., \j[k) U^} is linearly independent is by mathematical induction. By definition, the set [\]\l) Ur"} is linearly independent. Now, suppose that (for an arbitrary integer j between 1 and fc - 1, inclusive) the set {UJ1* Uri°, U(,2) Ur2),..., U\J) Urf} is linearly independent. Then, it suffices to show that the se"t (Uj" V«\ U<2) U«> U(/+1) U^} is linearly independent.
166 17. Intersections and Sums of Subspaces According to Part (a), Uj+i and U\ -\ h Uj are essentially disjoint. And, clearly, U^ uj}\ U(L2) l#2), ..., U^ U<f are in the subspace Wi + W2 + • • • + W/. Thus, it follows from Lemma 17.1.3 that the set {U^ U<|\ Uf2* U<2) U(/+1) 13^+0} is linearly independent. (e) Suppose that Wi, £^,..., £4 are independent. Let Ui, U2,..., Ujt represent nonnull matrices in U\, Ui,..., £4, respectively. Then, it follows from Part (d) that the set {Ui, U2 Ujt} is linearly independent. Conversely, suppose that, for every nonnull matrix Ui in U\% every nonnull matrix U2 in U2. •.., and every nonnull matrix Ujt in Wjt, Ui, U2 Ujt are linearly independent. \iU\,Ui Uk were not independent, then, for some nonempty subset [j\ jr) of the first k positive integers, there would exist nonnull matrices Uy,,..., Vjr in Uh,..., Ujr, respectively, such that 11,,+--- + 1^=0, and the set {Ui, U2 Ujt} (where, for j $ [ji yr}, Uy is an arbitrary nonnull matrix in Uj) would be linearly dependent, which would be contradictory. Thus, U\, Ui Uk are independent. (f) It is clear from the result of Exercise 5 that S spans U\-\ 1-¼. (1) Now, suppose that U\,..., Uk are independent. Then, it is evident from Part (d) that S is a linearly independent set. Thus, S is a basis for U\ -\ h Uk. (2) Alternatively, suppose that U\,..., Uk are not independent. Then, for some (nonempty) subset [jr,..., jr] of the first k positive integers, there exist nonnull matrices U/,,.. Further, for m = ., Uyr in Ujl Ujr, respectively Uy,+■•■ + %=». = 1,...,1-, Pjm 1=1 where c{m) cjj are scalars (not all of which can be zero) and Uf0 Uj£ are the matrices in S/m. Thus, r p/,„ m=l ;=i implying that 5 is a linearly dependent set. We conclude that S itself is not a basis and consequently (in light of Theorem 4.3.11) that S contains a proper subset that is a basis for U\ -\ \-Uk- (g) Part (g) is an immediate consequence of Part (f). EXERCISE 7. Let Aj A* represent matrices having the same number of rows, and let Bi Bjt represent matrices having the same number of columns.
17. Intersections and Sums of Subspaces 167 Adopting the terminology of Exercise 6, use Part (g) of that exercise to show (a) that if C(Ai) C(Ajt) are independent, then rank(Ai A*) = rank(Ai) + • • • + rank(Ajt), and if C(Ai) C(Ajt) are not independent, then rank(Ai A*) < rank(Ai) -\ h rank(Ajt) and (b) that if 7£(Bi) ft(Bjt) are independent, then /B,\ rank I : I = rank(Bi) -\ h rank(Bjt), W and if 7?.(Bi), ..., 7£(Bjt) are not independent, then /BA rank : <rank(Bi) + —|-rank(B/t). W Solution, (a) Clearly, rank(Ai) + • • • + rank(Ajt) = dim[C(Ai)] + • • + dim[C(Ajt)]. And, in light of equality (1.6), rank(Ai A*) = dim[C(A,..., A*)] = dim[C(Aj) + • • • + C(Ajt)]. Thus, it follows from Part (g) of Exercise 6 that if C(Ai) C(Ajt) are independent, then rank(Ai,..., A*) = rank(Ai) H h rank(Ajt), and if C(Ai),..., C(Ajt) are not independent, then rank(Ai A*) < rank(Ai) -\ 1- rank(Ajt). (b) The proof of Part (b) is analogous to that of Part (a). EXERCISE 8. Letting A represent an m x n matrix and B an m x p matrix, show, by for instance using the result of Part (c)-(2) of Exercise 4 in combination with the result C(A, B) = C[A, (I - AA")B] = C(A) 0 C[(I - AA")B], (*) that (a) C[(I - AA")B] = C(I - AA") n C(A, B) and
168 17. Intersections and Sums of Subspaces (b) C[(I - PA)B] = M{A') n C(A, B). Solution, (a) According to result (*) [or, equivalently, the first part of Corollary 17.2.9], C(A, B) = C(A) + C[(l - AA")B]. Thus, observing that C[(I - AA~)B] C C(I - AA~) and making use of the result of Part (c)-(2) of Exercise 4 (and also of Lemma 17.2.7), we find that C(I - AA") fl C(A. B) = C{\ - AA") n [C[(I - AA")B] + C(A)} = C[(I - AA")B] + [C(I - AA") fl C(A)] = C[(I-AA-)B] + {0} = C[(I-AA")B]. (b) According to Part (1) of Theorem 12.3.4, (A'A)~A' is a generalized inverse of A. Substituting this generalized inverse for A~ in the result of Part (a) and making use of Lemma 12.5.2, we find that C[(I - PA)B] = C{\ - PA) n C(A. B) = A/r(A,)flC(A,B). EXERCISE 9. Let A = (T. U) and B = (V, 0), where T is an m x p matrix, U an //? x q matrix, and V an n x p matrix, and suppose that U is of full row rank. Show that 1Z{A) and ft(B) are essentially disjoint [even if ft(T) and ft(V) are not essentially disjoint]. Solution. Let x; represent an arbitrary [1 x (p + q)] vector in 1Z(A) fl 72(B). Then, x' = r'A and x' = s'B for some (row) vectors r' and s'. Partitioning x' as x' = (x7,, x^) (where x', is of dimensions 1 x p), we find that (x'l,'s:2) = r,(T,V) = (r,T,r,V) and similarly that (x;^2) = s'(v.o) = (s'v,o). Thus, r'U = xiy = 0, implying (since the rows of U are linearly independent) that r' = 0 and hence that x; = 0. We conclude that 72(A) and 72(B) are essentially disjoint [even if 72(T) and 72(V) are not essentially disjoint]. EXERCISE 10. To what extent does the formula rank CR ^J = rank(U) + rank(V) + rank[(I - UIT)T(I - V~V)] (*) [where T is an m x p matrix, U an m x q matrix, and V an n x /? matrix] simplify in (a) the special case where C(T) and C(V) are essentially disjoint [but 7v(T) and
17. Intersections and Sums of Subspaces 169 Tl(\) are not necessarily essentially disjoint] and (b) the special case where K{T) and 7£(V) are essentially disjoint. Solution, (a) If C(T) and C(U) are essentially disjoint, then {since C[T(I - V" V)] C C(T)} C[T(I-V~ V)] andC(U) are essentially disjoint, and (in light of Corollary 17.2.10) formula (*) [or, equivalently, formula (2.15)] simplifies to ■PS)- rank (y Q J = rank(U) + rank(V) + rank[T(I - V~V)]. (b) If 7£(T) and 71(V) are essentially disjoint, then it follows from an analogous line of reasoning that formula (*) [or, equivalently, formula (2.15)] simplifies to *(* 9- rank ( y J = rank(U) + rank(V) + rank[(I - UlT )T]. EXERCISE 11. Let T represent an /» x p matrix, U an m x q matrix, and V annxp matrix. Further, define Er = I - TT~, Ft = I - T~T, X = E7U, (T— — T—UX—Et\ _-__- J is a generalized inverse of the partitioned matrix (T, U) and (b) that the partitioned matrix (T~— F7-Y~VT~, F^Y") is a generalized inverse of the partitioned matrix I v). Do so by applying formula (E.1) from Part (a) of Exercise 10.10 to the (T U\ /T 0\ ft J and i v ft J and by making use of the result that for any generalized inverse G = I ' I of the partitioned matrix (A, B) and any generalized inverse H = (Hj, H?) of the partitioned matrix ( r j (where A is an m x n matrix, B an wi x p matrix, and Ca^x« matrix and where G\ has n rows and Hj m columns), (1) Gj is a generalized inverse of A and G2 a generalized inverse of B if and only if C( A) and C(B) are essentially disjoint, and, similarly, (2) Hj is a generalized inverse of A and Ho a generalized inverse of C if and only if 11(A) andft(C) are essentially disjoint. Solution, (a) Upon setting V = 0 and W = 0 (in which case Y = 0, Q = 0, and Z = 0) and choosing Y~ = 0 and Z~ = 0 in formula (E.1) [from Part (a) of /T U\ Exercise 10.10], we obtain as a generalized inverse for I ft ft I the partitioned matrix /T--T-UX-ET 0\ G={ X-ET 0| We conclude, on the basis of the cited result (or, equivalently, Theorem 17.3.3). (T— — T—UX—Et\ _____ ) is a generalized inverse of (T, U).
170 17. Intersections and Sums of Subspaces (b) Upon setting U = 0 and W = 0 (in which case X = 0, Q = 0, and Z = 0) and choosing X~ = 0 and Z~ = 0 in formula (E.1) [from Part (a) of Exercise (T 0\ v ft ) the partitioned matrix -C F7Y-VT- FrY~ 0 0 We conclude, on the basis of the cited result (or, equivalently, Theorem 17.3.3), that (T~ - FrY~VT~, FTY~) is a generalized inverse of (y J. EXERCISE 12. Let T represent anmxp matrix, U an m x q matrix, and V an n x p matrix. And, let I " *2 J (where Gn is of dimensions p x m) (T U\ v ft I. Show that (a) if G11 is a generalized inverse of T and Gj 2 a generalized inverse of V, then 7£(T) and 11(V) are essentially disjoint, and (b) if Gj t is a generalized inverse of T and G21 a generalized inverse of U, then C(T) and C(U) are essentially disjoint. Solution. Clearly, /TGnT + UG21T + TG12V + UG22V TGj ,U + UG21IA V VG11T + VG12V VGnU ) = (v o){g2\ G22JVV oj = (v oj- (S1) (a) Result (S.l) implies in particular that VGnT = V-VGj2V. (S.2) Now, suppose that Gi j is a generalized inverse of T and G12 a generalized inverse of V. Then, equality (S.2) reduces to VGnT = 0, and it follows from Corollary 17.2.12 that ft(T) and 11(V) are essentially disjoint. (b) The proof of Part (b) is analogous to that of Part (a). EXERCISE 13. (a) Generalize the result that, for any two subspaces U and V of 1Zmxn, dim(U + V) = dim(W) + dim(V) - dim(U n V), (*) Do so by showing that, for any k subspaces U\,...Mk> dim(W| + • • • +Uk) = dim(Wj) + • • • + dim(24) Jfc - J^dim[(W| + • • • + Ui-{) nHi]. (E.2)
17. Intersections and Sums of Subspaces 171 (b) Generalize the result that, for any m x n matrix A, m x p matrix B, and qxn matrixC, rank(A, B) = rank(A) + rank(B) - dim[C(A) n C(B)], rank \t J = rank(A) + rank(C) - dim[ft(A) n 11(C)]. Do so by showing that, for any matrices Ai A* having the same number of rows, rank(Ai Ajt) = rank(Aj) -\ h rank(Ajt) Jt -J]dim[C(A,,...Al_i)nC(Ai)] /=2 and, for any matrices B\,..., B* having the same number of columns, /B, \ Q- rankl : | = rank(Bi) + ••• + rank(B/t) - £dim[ft| : I nR(B,)]. ta-i/ Solution, (a) The proof is by mathematical induction. In the special case where k = 2, equality (E.2) reduces to the equality dim(Wj + ½) = dim(Wi) +dim(W2) - dim(U\ nU2), which is equivalent to equality (*) and whose validity was established in Theorem 17.4.1. Suppose now that equality (E.2) is valid for k = k' (where k' > 2). Then, making use of Theorem 17.4.1, we find that dim(Wi+- --+^+^+0 = dim(Wj + • • • +lfa) + dim(24'+i) -dim[(Wj +--+Uk>)nUk'+i) k' = dim(Wi) + • • • + dim(^) - ^dim[(Wj + • • • + W,-i) nty] /=2 + dim(^+i) - dim[(Wi + ■ • • +W*0 nUk>+\] k'+i = dim(Wj) + • • • + dim(Uk>+i) - £] dim[(U\ + • • • + Ut-\) nty], i=2 thereby completing the induction argument.
172 17. Intersections and Sums of Subspaces (b) Applying Part (a) with U\ = C(A\) Uk = C(Ajt) [and recalling result (1.6)], we find that rank(Aj Ajt) = dim[C(Aj,... A*)] = dim[C(Aj) + -..+C(Ajt)] = dim[C(Aj )] + ••• + dim[C(A*)] Jt -^dim{[C(Aj) + ...+C(A;_i)]nC(A/)} /=2 = rank(Aj) -\ h rank(Ajt) it -£dim[C(Ai A/_i)nC(A/)]. i=2 And, similarly, applying Part (a) with U\ = ft(Bj) Uk = ft(Bjt) [and recalling result (1.7)], we find that /B.\ /BA rank : = dim [ft : ] W W = dim[ft(Bi) + -..+ft(B*)] = dim[ft(Bi)] +•. + dim[ft(Bjt)] Jt - ]Tdim{[ft(Bi) + • • - + ft(B,-i)] fl ft(B,)} /B.\ = rank(Bi) + ---+rank(Bjfc)-^dimra : nft(B,-)]. i=2 W-./ EXERCISE 14. Show that, for any m x n matrix A, /? xq matrix C, and q x p matrix B, rank{[I - CB(CB)~]C[I - (AC)~AC]} = rank(A) + rank(C) - rank(AC) - n + rank{[I - CB(CBr](I - A"A)}. Hint. Apply the equality rank(AC) = rank(A) + rank(C) - » + rank[(I - CC~)(I - A~A)] (*) to the product A(CB), and make use of the equality rank(ACB) = rank(AC) + rank(CB) - rank(C) + rank ([I - CB(CB)~] C [I - (AC)~AC]}. (**)
17. Intersections and Sums of Subspaces 173 Solution. Making use of equality (*) [or equivalently equality (5.8)], we find that rank(ACB) = rank[A(CB)] = rank(A) + rank(CB) - n + rank{[I - CB(CB)~](I - A"A)}. (S.3) And upon equating expression (**) [or equivalently expression (5.5)] to expression (S.3), we find that rank{[I - CB(CB)~]C[I - (AC)"AC]} = rank(A) + rank(C) - rank(AC) - n + rank{[I - CB(CB)~](I - A~A)}. EXERCISE 15. Show that if annxn matrix A is the projection matrix for a subspaceWofft"xl alongasubspaceVofft"xl (where^eV = ft"xl),then A' is the projection matrix for Vx along UL [where UL and Vx are the orthogonal complements (with respect to the usual inner product and relative to TZ"xl)ofU and V, respectively]. Solution. Suppose that A is the projection matrix for li along V (where W0V = ft"xl). Then, according to Theorem 17.6.14, A is idempotent, U = C(A), and V = C(\ — A). And, since (according to Lemma 10.1.2) A' is idempotent, it follows from Theorem 17.6.14 that A' is the projection matrix for C(A') along N(A.'). Moreover, making use of Corollary 11.7.2 and of Lemma 12.5.2, we find that C(A') = M(l - A') = CHl - A) = V1 and that Af(\') = C1(\)=U±. EXERCISE 16. Show that, for any n x p matrix X, XX- is the projection matrix for C(X) along MXX"). Solution. According to Lemma 10.2.5, XX" is idempotent. Thus, it follows from Theorem 17.6.14 that XX" is the projection matrix forC(XX~) along J\f(XX~). Moreover, according to Lemma 9.3.7, C(XX~) = C(X). EXERCISE 17. Let Y represent a matrix in a linear space V of m x n matrices, and let U\ Uk represent subspaces of V. Adopting the terminology and using the results of Exercise 6, show that if U\ Uk are independent and if U\ -\ I-Z4 = V, then (a) there exist unique matrices Zj,..., Zjt in U\,..., t/jt, respectively, such that Y = Zi H h Zjt and (b) for / = 1 fc, Z,- equals the projection of Y onU{ alongU\ + • • • +Ut-\ +Ui+\ + • • • +24-
174 17. Intersections and Sums of Subspaces Solution. Suppose that U\,..., Uk are independent and that U\-\ h Z4 = V. (a) It follows from the very definition of a sum (of subspaces) that there exist matrices Zi,..., Zjt in U\,..., tfo, respectively, such that Y = Zi -\ l-Z*. For purposes of establishing the uniqueness of Zi Z*, let Z\ Z£ represent matrices (potentially different from Z\,..., Zjt) in U\ 24, respectively, such that Y = ZJ + • • • + ZJ. Then, (Z?-Zi) + ... + (Z£-Z*)=Y-Y = 0, and (fori = 1 *)Zf-Z/ eW/.Thus.Zf-Z,- =OandhenceZf = Z, (/ = 1 Ic), thereby establishing the uniqueness of Zi Z*. (b) That (for i = 1 k) Z,- equals the projection of Y on Ui along Wi H h U,-\ + W/+i -\ h i4 is evident upon observing that [as a consequence of Part (b) of Exercise 6] U-, and U\ -\ \-Ui-i +Ui+i -\ VUu are essentially disjoint and that Y-Zl-=Zi+...+Z/_i+Zf+i+...+Z*eWi+. --+^1-1+^1+1+---+½. EXERCISE 18. Let U and W represent essentially disjoint subspaces (of TZnxl) whose sum is7lnxl, and let U represent any n x s matrix such that C(U) = U and Wanynx/ matrix such that C(W) = W. (a) Show that the n x (s + /) partitioned matrix (U, W) has a right inverse. (b) Taking R to be an arbitrary right inverse of (U, W) and partitioning R as R = (p1 I (where Ri has s rows), show that the projection matrix for U along W equals URi and that the projection matrix for W along U equals WR2. Solution, (a) In light of result (1.4), we have that rank(U, W) = dim[C(U, W)] = dim(W + W) = dim(ft") = n. Thus, (U, W) is of full row rank, and it follows from Lemma 8.1.1 that (U, W) has a right inverse. (b) For j = 1,...,«, let ey- represent the 7th column of I,,; let z; represent the projection of e/ on U along W; let ry, ri;, and r2; represent the jth columns of R, Ri, and R2, respectively, and observe that r; = [ ly J. By definition, (U, W)R = I„, implying that (for j = 1 n) (U, W)r; = e;. Thus, it follows from Corollary 17.6.5 that (for j = 1 n) z; = Uriy. We conclude (on the basis of Theorem 17.6.9) that the projection matrix fort/ along W equals (21 z„) = (Ur,i Uri„)=URi. And, since URi + WR2 = I„, we further conclude (on the basis of Theorem 17.6.10) that the projection matrix for W along U equals I - URj = WR2.
17. Intersections and Sums of Subspaces 175 EXERCISE 19. Let A represent the (n xn) projection matrix for a subspace U offt"xl along a subspace V of TZnxl (whereWe V = ft"xl),letB represent the (;i x /i) projection matrix for a subspace VV of TZnx l along a subspace X of 1Z" x! (where VV © X = 7£"x *), and suppose that A and B commute (i.e., that BA= AB). (a) Show that AB is the projection matrix for U H VV along V + X. (b) Show that A + B - AB is the projection matrix for U + W along VC\X. [Hint for Part (b). Observe that I - (A + B - AB) = (I - A)(I - B), and make use of Part (a).] Solution, (a) According to Theorem 17.6.13, A and B are both idempotent, so that (AB)2 = A(BA)B = A(AB)B = A2B2 = AB. Thus, AB is idempotent, and it follows from Theorem 17.6.14 that AB is the projection matrix for C(AB) along jV(AB). It remains to show that C(AB) = U n VV and jV(AB) = V + X or equivalently (in light of Theorem 17.6.14) that C(AB) = C(A) H C(B) and AA(AB) = N{\) + Af(B). Clearly, C(AB) c C(A) and (since AB = BA) C(AB) C C(B), so that C(AB) C C(A) H C(B). And, for any vector y in C(A) fl C(B), it follows from Lemma 17.6.7 that y = Ay and y = By, implying that y = ABy and hence that y e C(AB). Thus, C(A) fl C(B) C C(AB), and hence [since C(AB) c C(A) fl C(B)] C(AB)=C(A)flC(B). Further, for any vector x in AT(A) and any vector y in A/"(B), AB(x + y) = ABx + ABy = B Ax + ABy = 0 + 0 = 0, implying that x + y e N{AB). Thus, N(A) + N(B) C ^(AB). And, for any vector z in N(AB) (i.e., any vector z such that ABz = 0), Bz e A/"(A), which since z = Bz+(I—B)z and since (I—B)z e AfflB) [as is evident from Theorem 11.7.1 or upon observing that B(I-B)z = (B-B2)z = 0] implies that z e Af(A)+N(JB).lt follows thatMAB) c AA(A)+^(8), and hence [since jV(A)+^(8) cAT(AB)] that AT(AB) = ^(A) +N(B)- (b) As a consequence of Theorem 17.6.10,1 — A is the projection matrix for V along U, and I — B is the projection matrix for X along W. Thus, it follows from Part (a) that (I - A) (I - B) is the projection matrix for V fl X along U + VV. Observing that A + B — AB = I — (I — A) (I — B), we conclude, on the basis of Theorem 17.6.10, that A + B - AB is the projection matrix for U + W along vnx. EXERCISE 20. Let V represent a linear space of n-dimensional column vectors, and let U and VV represent essentially disjoint subspaces whose sum is V. Then, an/ixn matrix A is said to be a projection matrix for U along VV if Ay is the projection of y on U along VV for every y e V — this represents an extension of the definition of a projection matrix for U along VV in the special case where V = 11". Further, let U represent an n x s matrix such that C(U) = U, and let W represent an«x/ matrix such that C(W) = VV.
176 17. Intersections and Sums of Subspaces (a) Show that an n x n matrix A is a projection matrix for U along W if and only if AU = U and AW = 0 or, equivalently, if and only if A' is a solution to the linear system I w, JB = I ft J (in an n x n matrix B). (b) Establish the existence of a projection matrix for U along W. (c) Show that if A is a projection matrix for U along W, then I—A is a projection matrix for W along U. (d) Let X represent any n x p matrix whose columns span N(W) or, equivalently, VVX. Show that an n x n matrix A is a projection matrix fort/ along W if and only if A' = XR* for some solution R* to the linear system U'XR = U' (in a p x n matrix R). Solution, (a) Clearly, an n x 1 vector y is in V if and only if y is expressible as y = Ub + Wc for some vectors b and c. Thus, an n x n matrix A is a projection matrix fort/ along W if and only if, for every (s x 1) vector b and every (t x 1) vector c, A(Ub+Wc) is the projection of Ub+Wc on U along W, or equivalently (in light of Corollary 17.6.2) if and only if, for every b and every c, A(Ub + Wc) = Ub. Now, if AU = U and AW = 0, then obviously A(Ub + Wc) = Ub for every b and every c. Conversely, suppose that A(Ub + Wc) = Ub for every b and every c. Then, A(Ub + Wc) = Ub for every b and for c = 0, or equivalently AUb = Ub for every b, implying (in light of Lemma 2.3.2) that AU = U. Similarly, A(Ub + Wc) = Ub for b = 0 and for every c, or equivalently AWc = 0 for every c, implying that AW = 0. (b) Clearly, the linear systems U'B = U' and W'B = 0 (in B) are both consistent. And, since (in light of Lemma 17.2.1) 7£(U') and 1Z(W) are essentially disjoint, we have, as a consequence of Theorem 17.3.2, that the combined linear system I , IB = I ft J is consistent. Thus, the existence of a projection matrix for U along W follows from Part (a). (c) Suppose that A is a projection matrix for U along W. Then, according to Part (a), AU = U and AW = 0. Thus, (I - A)W = W, and (I - A)U = 0. We conclude [on the basis of Part (a)l that I — A is a projection matrix for W along U. (d) In light of Part (a), it suffices to show that A' is a solution to the linear system [ w, JB = I . Win B) if and only if A' = XR* for some solution R* to the linear system U'XR = U'. Suppose that A' = XR* for some solution R* to U'XR = U'. Then, U'A' = U', and (since clearly W'X = 0) W'A; = 0. Thus, A' is a solution to (w, JB = ( fl ). Conversely, suppose that A' is a solution to j w, )B = ( ft ) or equivalently that U'A' = U' and W'A' = 0. Then, according to Lemma 11.4.1, C(A') c MW), or equivalently C(A') C C(X), and consequently A' = XR* for some matrix R*.
17. Intersections and Sums of Subspaces 177 And, U'XR* = U'A' = U\ so that R* is a solution to U'XR = U'. EXERCISE 21. Let Wj Uk represent independent subspaces of TZ"xl such thatt/j -\ 1-½ =TZnxl (where the independence of subspaces is as defined in Exercise 6). Further, letting j,- = dim(%-) (and supposing that .9, > 0), take U,- to be any n x Sj matrix such that C(U/) = Ui (i = 1 k). And, define /BA B = (Ui Uj)-1, partition B as B = : (where, for i = 1 k. B, has W Si rows), and let H = B'B or (more generally) let H represent any matrix of the form H = B', AjBj + B'2A2B2 + • • • + B^A^Bjt, (E.3) where A|, A2,..., A* are symmetric positive definite matrices. (a) Using the result of Part (g)-(l) of Exercise 6 (or otherwise), verify that the partitioned matrix (Uj,..., Ujt) is nonsingular (i.e., is square and of rank n). (b) Show that H is positive definite. (c) Show that (for j £ i = 1 k)Ui and Uj are orthogonal with respect to H. (d) Using the result of Part(a)-(2) of Exercise 3 (or otherwise), show that, for / = 1 fc, (1) U\ + • • • + Ui-\ + Uj+i -\ +Uk equals the orthogonal complement Uf- of Ui (where the orthogonality in the orthogonal complement is with respect to the bilinear form x'Hy) and (2) the projection of any n x 1 vector y on Ui alongU\ H YUi-\ + Ui+\ -\ \-Uk equals the orthogonal projection of y on Ui with respect to H. (e) Show that if, for j £ i = 1 A', Ui and Uj are orthogonal with respect to some symmetric positive definite matrix H*, then H* is expressible in the form (E.3). Solution, (a) Clearly, dim(Wi + • • • +Uk) = dim(7e"xl) = n. Thus, making use of Part (g)-(l) of Exercise 6, we find that si + • • • +sk = dim(Wi) + • • • + dim(Uk) = dim(Wi + • • • +Uk) = n. And, making use of result (1.6), we find that rank(Ui Ujt) = dim[C(Ui U*)] = dim[C(Ui ) + •••+ C(U*)] = dim(Wi+---+½) =«. (b) Clearly, H = B; diag(Ai A*)B. Thus, since (according to Lemma 14.8.3) diag(Ai AjO is positive definite, it follows from Corollary 14.2.10 that H is positive definite. (c) In light of Lemma 14.12.1, it suffices to show that (for ; £ i) U{HUy = 0.
178 17. Intersections and Sums of Subspaces By definition, /B,Uj BiU2 ... BiU*\ B2U, B2U2 ... B2U* VBAUi BjtU2 ... BjtUjt/ = B(U,,U2 U*) = I„ = (hx 0 ... 0 \ 0 I„ 0 V° ° V implying in particular that (for j # i) B;U, = 0 and (for r # j) BrUy = 0. Thus, for j 56 1, UjHUy = (ByUf )'AyByUy + £ UjB;.ArBrU,- = 0 + 0 = 0. (d) (1) According to Part (c), Ui is orthogonal to t/j Ui-\, t//+i £4. Thus, making repeated (fc - 2 times) use of Part (a)-(2) of Exercise 3, we find that Ui is orthogonal to U\ -\ Vhk-\ +W/+1 -\ 1-½. We conclude (on the basis of Lemma 17.7.2) that U\ + • • • + W/-1 +Ui+i + • • • +Uk = U±. (2) That the projection of y on Ui along U\ + • • • + Ut-\ + Ul+\ + • • • + Uk equals the orthogonal projection of y on Ui (with respect to H) is [in light of Part (1)] evident from Theorem 17.6.6. (e) Suppose that, for j ^ i = 1 kMi and Uj are orthogonal with respect to some symmetric positive definite matrix H*. Then, according to Corollary 14.3.13, there exists an n x n nonsingular matrix P such that H* = P'P. Further, P = PI„=P(U, Ujt) : =L|B|+---+LitBib. where (for / = 1 k) L/ = PU/. And, making use of Lemma 14.12.1. we find that,for; #1 = 1 k, LjLj = UfP'PU, = UjHUU, = 0. Thus, H* = (L1B1 + • ■ ■ + LftBt)'(LiB| + • • • + UBk) = B'^LjBi + B'2L'2L2B2 + • • • + B^LfBt = B; AiBj + B'2A2B2 + • • • + BiAftB*. where (for/ = 1,..., k)A{ = LjL/ [which is a symmetric positive definite matrix, as is evident from Corollary 14.2.14 upon observing that rank(L,) = rank(PU,) = rank(U/) =5,-].
18 Sums (and Differences) of Matrices EXERCISE 1. Let R represent an n x n matrix, S an n xm matrix, T an m x m matrix, and V anm x n matrix. Derive (for the special case where R and T are nonsingular), the formula |R + STU| = |R| IT + TUR-^TI/m. Do so by making two applications of the formula UlTMW-VT-'UI. (*) T U V w W V U T — one with W set equal (in which V is an n x m matrix and Wannxn matrix and in which T is assumed I R —STl to be nonsingular) to the partitioned matrix __. to R, and the other with T set equal to R. Solution. Suppose that R and T are nonsingular. Then, making use of formula (*) (or equivalently the formula of Theorem 13.3.8), we find that R -ST |TU and also that = |T| |R - (-ST)T_1TU| = |T| |R + STU| |R TU -ST T Thus, = |R| |T- (TU)R_1(-ST)| = |R| |T + TUR-1ST|. |T| |R + STU| = |R| IT + TUR-'STI,
180 18. Sums (and Differences) of Matrices or equivalently IR + STUI = |R| |T + TUR"IST|/|T|. EXERCISE 2. Let R represent an n x n matrix, S an n x m matrix, T an m x m matrix, and U an m x n matrix. Show that if R is nonsingular, then |R + STU| = |R| |IW +UR-!ST| = |R| |I,H +TUR~1S|. Do so by using the formula |R + STU| = |R| |T| |T_1 + UR_1S|, (*) (in which R and T are assumed to be nonsingular), or alternatively the formula |IH + SU| = |I,H + US| or the formula |R + STU| = |R| |T + TUR-'STI/m. Solution. Note that R + STU = R + (ST)IIHU, (S.l) R + STU = R + SI,„ (TU). (S.2) Now, suppose that R is nonsingular. By applying formula (*) (or equivalently the formula of Theorem 18.1.1) to the right side of equality (S.l) [i.e., by applying formula (*) with ST and lm in place of S and T, respectively], we find that |R + STU| = |R| |IW| II"1 +UR-'ST| = |R| |I/H +UR~1ST|. Similarly, by applying formula (*) to the right side of equality (S.2) [i.e., by applying formula (*) with I„, and TU in place of T and U, respectively], we find that |R + STU| = |R| |I,„| II"1 +TUR~1S| = |R| |I,„ +TUR~1S|. EXERCISE 3. Let A represent an n x n symmetric nonnegative definite matrix. Show that if I — A is nonnegative definite and if |A| = 1, then A = I. Solution. Suppose that I - A is nonnegative definite and that |A| = 1. Then, as a consequence of Corollary 14.3.12. A is positive definite. And, HI = 1 = |A|. Thus, it follows from Corollary 18.1.7 (specifically from the special case of Corollary 18.1.7 where C = I) that I = A. EXERCISE 4. Show that, for any n x n symmetric nonnegative definite matrix B and for any n x n symmetric matrix C such that C — B is nonnegative definite, |C|>|C-B|.
18. Sums (and Differences) of Matrices 181 with equality holding if and only if C is singular or B = 0. Solution. Let A = C - B. Then, C - A = B. So, by definition, A is a (symmetric) nonnegative definite matrix, and C - A is nonnegative definite. Thus, it follows from Corollary 18.1.8 that |C|>|C-B|, with equality holding if and only if C is singular or C = C - B, or equivalently if and only if C is singular or B = 0. EXERCISE 5. Let A represent a symmetric nonnegative definite matrix that has been partitioned as where T is of dimensions m x m and W of dimensions n x n (and where U is of dimensions m x n). And, define Q = W—U'T~U (which is the Schur complement ofT). (a) Using the result that the symmetry and nonnegative definiteness of A imply the nonnegative definiteness of Q and the result of Exercise 14.33 (or otherwise), show that |W| > |U T~U|, with equality holding if and only if W is singular or Q = 0. (b) Suppose that n = m and that T is nonsingular. Show that |W| |T| > |U|2, with equality holding if and only if W is singular or rank(A) = m. (c) Suppose that n = m and that A is positive definite. Show that |W| |T| > |U|2. Solution, (a) According to the result of Exercise 14.33, UT~U is symmetric and nonnegative definite. Further, W is symmetric. And, in light of the result that the symmetry and nonnegative definiteness of A imply the nonnegative definiteness of Q - W - U'T~U [a result that is implicit in Parts (1) and (2) of Theorem 14.8.4], it follows from Corollary 18.1.8 that |W| > |U T~U|, with equality holding if and only if W is singular or W = UT~U, or equivalently if and only if W is singular or Q = 0. (b) Since (in light of Corollary 14.2.12) |T| > 0 and since lUX-'UI = lU'l IX-1! |U| = |U|2/|T|,
182 18. Sums (and Differences) of Matrices |W| |T| > |U|2 & |W| > lUT-'UI and |W| |T| = |U|2 & |W| = lU'T-'UI. Moreover, in light of Theorem 8.5.10, rank(A)=m <s> rank(Q) = 0 «£> Q = 0. Thus, it follows from Part (a) that |W| |T| > |U|2, with equality holding if and only if W is singular or rank(A) = m. (c) We have (in light of Lemma 14.2.8 and Corollary 14.2.12) that rank(A) = 2m > m and that W (and T) are nonsingular. Thus, it follows from Part (b) that |W| |T| > |U|2. EXERCISE 6. Show that, for any n x p matrix X and any symmetric positive definite matrix W, |X'WX| IX'W'XI > |XX|2 . (E.l) [Hint. Begin by^showing that the matrices X'X(X'WX)~X'X and XW'X - X X(X WX)~X X are symmetric and nonnegative definite.] Solution. Let A = X'X(X'WX)-X'X and C = X'w'X. Then, making use of Part (6') of Theorem 14.12.11, we find that A = XW-^Px.wW-'X = x'w-1Px,wWPx,wW-1X = (Px.wW-1X)'W(Px.wW-1X), so that (in light of Theorem 14.2.9) A is symmetric and nonnegative definite. Further, C is symmetric, and, making use of Part (9) of Theorem 14.12.11, we find that C - A = XW"1 W(I - PxavJW-'X = XW-'d - Px.w)'W(I - PxavJW-'X = [(I - Px.wJW-'Xj'WKI - Px.wJW-^], so that C - A is nonnegative definite. Thus, it follows from Corollary 18.1.8 that ix'w-'xi > ix'x(x wxrx'xi. (S.3) If rank(X) = /?, then (in light of Theorem 14.2.9 and Lemma 14.9.1) |XWX| > 0 and (in light of Theorems 13.3.4 and 13.3.7) ix'x(x wxrx'xi = ix'xi Kx'wxr'i ix'xi = ix'xi2/ix'wxi.
18. Sums (and Differences) of Matrices 183 in which case inequality (S3) is equivalent to inequality (S.l). Alternatively, if rank(X) < p, then [since (according to Corollary 14.11.3) rank(XWX) = rank(X) and rank(XX) = rank(X)] both sides of inequality (E.1) equal 0 and hence inequality (E.1) holds as an equality. EXERCISE 7. (a) Show that, for any n x n skew-symmetric matrix C, Hi + C|>1. with equality holding if and only if C = 0. (b) Generalize the result of Part (a) by showing that, for any n x n symmetric positive definite matrix A and any n x n skew-symmetric matrix B, |A + B|>|A|, with equality holding if and only if B = 0. Solution, (a) Clearly, |I + C| = |(I + C)'| = |I + C'| = |I-C|. so that li+ci2 = |i+ci n - ci = ici+o(i - oi = |i - cci = |i+c'q. Moreover, since C C is symmetric and nonnegative definite, it follows from Theorem 18.1.6 that |I + C'C|>|I|, with equality holding if and only if C C = 0 or equivalently if and only if C = 0. Since |I| = 1, we conclude that H + C|2>1, with equality holding if and only if C = 0. To complete the proof, it suffices to show that |I+C| > 0. According to Lemma 14.6.4, C is nonnegative definite. Thus, we have (in light of Lemma 14.2.4) that 1+C is positive definite and hence (in light of Corollary 14.9.4) that |I + C| > 0. (b) According to Corollary 14.3.13, there exists a nonsingular matrix P such that A = PP. Then, A + B = P(I + C)P, where C = (P_1 )'BP_1. Moreover, since (according to Lemma 14.6.2) C is skew- symmetric, we have [as a consequence of Part (a)] that H + C|>1, with equality holding if and only if C = 0 orequivalently if and only if B = 0. The proof is complete upon observing that, since |A+B| = |P|2|I+C| and |A| = |P|2
184 18. Sums (and Differences) of Matrices (and since |P| # 0), |A + B| > |A| <s> |I + C| > 1, and |A + B| = |A| & |I + C| = 1. EXERCISE 8. (a) Let R represent an n x n nonsingular matrix, and let B represent an n x n matrix of rank one. Show that R + B is nonsingular if and only if tr(R_1B) # — 1, in which case (R + B)-1 =R_1 -[1 +tr(R-,B)r,R~,BR-1. (b) To what does the result of Part (a) simplify in the special case where R = I„? Solution, (a) It follows from Theorem 4.4.8 that there exist n-dimensional column vectors s and u such that B = su'. Then, as a consequence of Corollary 18.2.10, we find that R + B is nonsingular if and only if uR"'s #-1- Moreover, upon applying result (5.2.6) (with b = u and a = R_1s), we obtain u'R_Is = trfR-W) = tr(R_,B). Thus, R+B is nonsingular if and only if tr(R_1B) # -1. And,iftr(R_1B) # -1, then we have, as a further consequence of Corollary 18.2.10, that (R + B)-1 =R-I-(l+u'R-1s)R-,su'R-1 = R-' -[1 +tr(R-1B)]_1R-1BR-1. (b) In the special case where R = I„, the result of Part (a) can be restated as follows: I,, + B is nonsingular if and only if tr(B) # — 1, in which case (Iw + B)-1 = I„ - [1 + tr(B)]-!B. EXERCISE 9. Let R represent an n x n matrix, S an n x m matrix, T an m x m matrix, and U an m x n matrix. Suppose that R is nonsingular. Show (a) that R+STU is nonsingular if and only if I„, + UR~l ST is nonsingular, in which case (R + STUr1 =R-' -R-^Td^+UR-^Tr^R-1. and (b) that R + STU is nonsingular if and only if I„, + TUR_1S is nonsingular, in which case (R + STUr1 =R_1 -R^SOU+TUR-'Sr'TUR-1. Do so by using the result—applicable when T (as well as R) is nonsingular—that R+STU is nonsingular if and only if T_1 +UR-1S is nonsingular, orequivalently if and only if T + TUR_1ST is nonsingular. in which case (R + STU)"1 = R"1 -R-'StT"1 +UR"1S)"1UR-1 = R"1 -R-'STCT + TUR-'STr'TUR'1.
18. Sums (and Differences) of Matrices 185 [Hint. Reexpress R + STU as R + STU = R + (ST)ImU and as R + STU = R + SI,„TU.] Solution, (a) Reexpress R + STU as R + STU = R+(ST)I„2U. Then, applying the cited result (or equivalently Theorem 18.2.8) with ST and lm in place of S and T, respectively, we find that R + STU is nonsingular if and only if lm + UR_1ST is nonsingular, in which case (R + STU)-1 = R-1 - R_1ST(IW + UR^ST^UR-1. (b) Reexpress R + STU as R + STU = R + SIW(TU). Then, applying the cited result (or equivalently Theorem 18.2.8) with l,„ and TU in place of T and U, respectively, we find that R + STU is nonsingular if and only if 1,, + TUR_1S is nonsingular, in which case (R + STU)-1 = R-1 - R_1S(Im + TUR-^S^TUR-1. EXERCISE 10. Let R represent an n x q matrix, San/ixm matrix, T an m x p matrix, and U a p x q matrix. Extend the results of Exercise 9 by showing that if ft(STU) C ft(R) and C(STU) c C(R), then the matrix R~ - R~ST(Ip + UR"ST)-UR- and the matrix R" - R-SOm + TUR_S)-TUR- are both generalized inverses of the matrix R + STU. Solution. Observe that R + STU can be reexpressed as R + STU = R+(ST)IPU and also as R + STU = R + SI,W(TU). Suppose now that ft(STU) C ft(R) and C(STU) C C(R). Then, upon applying Theorem 18.2.14 with ST and Ip in place of S and T, respectively, we find that R" - R~ST(Ip + UR-ST)"UR- is a generalized inverse of the matrix R + STU. And, upon applying Theorem 18.2.14 with I„, and TU in place of T and U, respectively, we find that R~ - R-SOn + TUR"S)-TUR-
186 18. Sums (and Differences) of Matrices is also a generalized inverse of R + STU. EXERCISE 11. Let R represent an n x q matrix, S an n x m matrix, T an m x p matrix, and U a p x q matrix. (' R —ST\ TIJ T )' and partition GasG=(rn r I2 J (where Gi i is of dimensions q x «). Show that Gj i is a generalized inverse of the matrix R + STU. Do so by using the result that, for any partitioned matrix A = ( " .12 j such that C(A2i) C C{\n) and 7£(Aj2) C 7£(A22> and for any generalized inverse I '* ~12 I of A (where Cii is of the same dimensions as A7,,), Cn is a generalized inverse of the matrix An — A12A70A21. (b)LetE/? = I-RR-,F/? = I-R~R,X = E/?ST, Y = TUF*,Ey = I-YY", Fx = I - X~X, Q = T + TUR-ST, Z = EyQF*, and Q* = FXZ~EY. Use the result of Part (a) of Exercise 10.10 to show that the matrix R~ - R~STQ*TUR- - R~ST(I - Q*Q)X-E/? - F/?Y-(I - QQ*)TUR- + F*Y-(I - QQ^QX-E/? (E.2) is a generalized inverse of the matrix R + STU. (c) Show that if ft(TU) C ft(R) and C(ST) C OR), then the formula R~ - R-STQ-TUR- (*) for a generalized inverse of R + STU can be obtained as a special case of formula (E.2). Solution, (a) It follows from the cited result (or equivalently from the second part of Theorem 9.6.5) that Gj i is a generalized inverse of the matrix R - (-ST)T-TU = R + STT-TU = R + STU. ( R — ST\ TIT T ) °btame(* by applying formula (10.E. 1). Partition GasG=[" ^12) (where Gj j is of dimensions q x /*), and assume that [in applying formula (10.E.1)] the generalized inverse of -X is set equal to -X~ [in which case F.v = I - (-X~ )(-X)]. Then, Gi i equals the matrix (E.2). and we conclude on the basis of Part (a) (of the current exercise) that the matrix (E.2) is a generalized inverse of the matrix R + STU. (c) Suppose that ft(TU) C ft(R) and C(ST) c C{R). Then, it follows from Lemma 9.3.5 that X = 0 and Y = 0 (so that F.y = I and E>- = I and consequently
18. Sums (and Differences) of Matrices 187 Q* is an arbitrary generalized inverse of Q). Thus, formula (*) [or equivalently formula (2.27)] can be obtained as a special case of formula (E.2) by setting X~ = 0 andY~ = 0. EXERCISE 12. Let Ai, Ao represent a sequence of m x n matrices, and let A represent another m x n matrix. (a) Using the result of Exercise 6.1 (i.e., the triangle inequality), show that if HAjt-AH-CUhenllAjfcll -* ||A||. (b) Show that if A* -> A, then ||Ajt|| ->• ||A|| (where the norms are the usual norms). Solution, (a) Making use of the triangle inequality, we find that IIAjtll = |(A* - A) + A|| < |A* - A|| + ||A|| and that |A|| = |At - (A* - A)|| < IIAjtH + |Ajt - A||. Thus, l|Ajtl|-||A||<||AA.-A||, and -(IIAjtH - UAH) = ||A|| - ||A*| < ||A*- A||, implying that I IIA*II - IIA|| | < ||A*-A||. Suppose now that || Ajt — A|| -> 0. Then, corresponding to each positive scaler €, there exists a positive integer p such that, for k > p, || Ajt — A|| < € and hence such that, for k > p, | ||Ajt|| - ||A|| | < €. We conclude that ||Ajt|| -> ||A||. (b) In light of Lemma 18.2.20, Part (b) follows from Part (a). EXERCISE 13. Let A represent an n x n matrix. Using the results of Exercise 6.1andofPart(b)ofExercisel2,showthatif||A|| < 1, then (fork = 0,1,2,...) ||(I-Ar1-(H-A + A2 + --. + AA')||<||Af+I/(l-||A||) (where the norms are the usual norms). (Note. If ||A|| < 1, then I — A is nonsin- gular.) Solution. Suppose that ||A|| < 1, and (for p = 0,1,2,...) let Sp = E,n=o A"' (where A0 = I). Then, as a consequence of Theorems 18.2.16 and 18.2.19, we have that(I -A)-1 = limp_co Sp, implying that (I - A)"1 - S* = ( lim Sp) - Sjt = lim (S/; - S*) p-*oo r p-*oo ' = lim Y A'",
188 18. Sums (and Differences) of Matrices and it follows from the result of Part (b) of Exercise 12 that lia —A>-"—Sjtll = Hm || T A"'||. (S.4) Moreover, making repeated use of the result of Exercise 6.1 (i.e., of the triangle inequality) and of Lemma 18.2.21, we find that (for p > k + 1) || J2 A"'ll^ E «AM«^ E 1*1" = IAI*1 E «A«"' (S'5> m=Jt+l m=k+\ m=k+\ w=0 It follows from a basic result on geometric series [which is example 34.8(c) in Bartle's (1976) book] that £~=0 ||A||W = 1/(1 - ||A||). Thus, combining result (S.5) with result (S.4), we find that p-k-\ ll(I-A)-1-SJt|| < lim [||A||*+I T ||A|H p-*°° to = iiai™ f; iad* = iai*+7u - iai>. m=0 EXERCISE 14. Let A and B represent /2 x n matrices. Suppose that B is nonsingular, and define F=B_1 A. Using the result of Exercise 13, show that if ||F|| < 1, then (for ^ = 0,1,2,...) H(B-A)-1 -(B"1 +FB"1 +F2B-! +-.-+FfcB"1)|| < IIB-^I l|F||*+7(l H|F||) (where the norms are the usual norms). (Note. If ||F|| < 1, then B — A is nonsin- gular.) Solution. Suppose that ||F|| < 1. Then, since B - A = B(I - F) (and since B - A is nonsingular), I — F is nonsingular, and (B- A)"1 = (1-F^B"1. Thus, making use of Lemma 18.2.21 and the result of Exercise 13, we find that ||(B - A)"1 - (B_1 + FB"1 + F2B-1 + • • • H-F^B"1)!! = ||[(I - F)"1 - (1 + F + F2 + • • • + F^B"11| < ||(I - F)-1 - (1 + F + F2 + • • • + F*)|| HE"11| < DB-Ml |F|*+I/(1 - |F|). EXERCISE 15. Let A represent an n x n symmetric nonnegative definite matrix, and let B represent an n x n matrix. Show that if B — A is nonnegative definite (in which case B is also nonnegative definite), then 1Z(A) C 7v(B) and C(A) c C(B).
18. Sums (and Differences) of Matrices 189 Solution. Define C = (I - B~B)'(B - A)(I - B~B). Since A is symmetric and nonnegative definite, there exists a matrix R such that A = R'R. Clearly, C = -(I - B"B)'A(I - B"B) = -[R(I - B~B)]'R(I - B~B). (S.6) Suppose now that B — A is nonnegative definite. Then, according to Theorem 14.2.9, C is nonnegative definite. Moreover, it is clear from expression (S.6) that C is nonpositive definite and symmetric. Consequently, it follows from Lemma 14.2.2 that C = 0 or equivalently that [R(I - B_B]'R(I - B~B) = 0, implying (in light of Corollary 5.3.2) that R(I - B~B) = 0 and hence (since A = R'R) that A(I - B~B) = 0. We conclude (in light of Lemma 9.3.5) that 11(A) C 11(B). Further, since B -- A is nonnegative definite, (B - A)' = B' - A' is also nonnegative definite. Thus, by employing an argument analogous to that employed in establishing that 11(A) C 11(B), it can be shown that ft(A') C ft(B') or equivalently (in light of Corollary 4.2.5) that C(A) C C(B). An alternative solution to Exercise 15 can be obtained by making use of Corollary 12.5.6. Suppose that B — A is nonnegative definite. And, let x represent an arbitrary vector in (^(B). Then, 0 < x'(B - A)x = -x'Ax < 0, implying that x'Ax = 0 and hence (in light of Corollary 14.3.11) that A'x = Ax = 0 or equivalently that x e CJ-(A). Thus, CL(B) C C"L(A), and it follows from Corollary 12.5.6 that C(A) C C(B). That 11(A) C 11(B) can be established via an analogous argument. EXERCISE 16. Let A represent annxn symmetric idempotent matrix, and let B represent an n x n symmetric nonnegative definite matrix. Show that if I — A — B is nonnegative definite, then BA = AB = 0. (Hint. Show that A'(I - A - B)A = —A'BA, and then consider the implications of this equality.) Solution. Clearly, A'(I - A - B)A = A'(A - A2 - BA) = A'(A - A - BA) = -A'BA. (S.7) Suppose now that I - A — B is nonnegative definite. Then, as a consequence of Theorem 14.2.9, A'(I - A - B)A is nonnegative definite, in which case it follows from result (S.7) that A'BA is nonpositive definite. Moreover, as a further consequence of Theorem 14.2.9, A'BA is nonnegative definite. Thus, in light of Lemma 14.2.2, we have that A'BA = 0. (S.8) And, since B is symmetric as well as nonnegative definite, we conclude (on the basis of Corollary 14.3.11) that BA = 0 and also [upon observing that AB = A,B/ = (BA),]thatAB = 0. EXERCISE 17. Let Ai Ajt represent n x n symmetric matrices, and define A = Ai + • • • + A*. Suppose that A is idempotent. Suppose further that
190 18. Sums (and Differences) of Matrices Ai Ajt_i are idempotent and that A* is nonnegative definite. Using the result of Exercise 16 (or otherwise), show that A/A; = 0 (for j ^ / = 1 k), that A* is idempotent, and that rank(Aft) = rank(A) — J^/r/ rank(A,-). Solution. Let Ao = I - A. Then, £f=0 A,- = I. Further, Ao (like Ai Ajt_i) is symmetric and idempotent, and (in light of Lemma 14.2.17) Ao, Ai Ajt_i (like Ajt) are nonnegative definite. Thus, for i = 1 k — 1 and j = i + 1,..., kt A,- is idempotent, Ay is nonnegative definite, and (since I - A,- - Ay = £*,=0 (m&j) Am) I - A,- - Ay- is nonnegative definite. And, it follows from the result of Exercise 16 that (for i = 1 k — 1 and / = / + 1 k) A,-Ay = 0 and Ay A/ = 0 or equiv- alently that, for j ^ / = 1 fc, A,-Ay = 0. Moreover, since Ai Aft are symmetric, we conclude from Theorem 18.4.1 that Aft (like Ai Aft_i) is idempotent and that ^=1 rank(A,) = rank(A) or, equivalently, rank(Aft) = rank(A) - 5^1,1 rank(A,-). EXERCISE 18. Let Ai Aft represent n x n symmetric matrices, and define A = Ai -\ h Aft. Suppose that A is idempotent. Show that if Aj Aft are nonnegative definite and if tr(A) < £f=l tr(A?), then A,-Ay = 0 (for j ^ / = 1 k) and Ai Aft are idempotent. Hint. Show that 52,-^,- tr(A,Ay) < 0 and then make use of the result that, for any two symmetric nonnegative definite matrices B and C (of the same order), tr(BC) > 0, with equality holding if and onlyifBC = 0. Solution. Clearly, A = A2 = (^ = ^+^,, i i i,j& so that tr(A) = *(£ A? + £ A/Ay) = £>(A?) + £ ^'^ and hence £ tr(A/Ay) = tr(A) - £ tr(A?). (S.9) '*./¥' i Suppose now that Aj,..., Aft are nonnegative definite and also that tr(A) < Y,i tr(A?). Then, it follows from result (S.9) that £tr(A,Ay)<0. And, since (according to Corollary 14.7.7, which is the result cited in the hint) tr(AjAy) > 0 (for all / and j £ /), we have that tr(A,Ay) = 0 (for all / and j # i). We conclude (on the basis of Corollary 14.7.7) that A,-Ay = 0 (for all i
18. Sums (and Differences) of Matrices 191 and ; # i). And, in light of Theorem 18.4.1 (and the symmetry of Ai A*), we farther conclude that A|,..., A* are idempotent. EXERCISE 19. Let A i A* represent n x n symmetric matrices such that Ai + • • • + A* = I. Show that if rank(Aj) + • • • + rank(A/t) = /i, then, for any (strictly) positive scalars ci,..., cjt, the matrix c\A\-\ h a A* is positive definite. Solution. Suppose that rank(Ai) -\ h rank(Ajt) = n. Then, it follows from Theorem 18.4.5 that Ai A* are idempotent. Thus, ciAi + ... + c*Ajt=l \y/ckAk/ \v/cjtAjt/ implying (in light of Corollary 14.2.14) that ciAj + h cjtAjt is nonnegative definite and (in light of Corollaries 7.4.5 and 4.5.6) that /V^A, rank(ciAi -\ \- C&A*) = rank I \VqAj \A*/ \Ak) = rank(A/lAi+..-+A^) = rank(Ai+--- + Ajt) = rank(I„) We conclude (on the basis of Corollary 14.3.12) that c\ Ai -\ h a A* is positive definite. EXERCISE 20. Let Ai A* represent n x n symmetric idempotent matrices such that A/Ay = 0 for j £ i = 1 k. Show that, for any (strictly) positive scalar cq and any nonnegative scalars ci,..., cjt, the matrix col + 5Z/=i C»A/ is positive definite (and hence nonsingular), and k k (c0I + J^ctAi)-1 = d0I + Y,diAi, i=i /=1 where do = l/c0 and (for / = 1 k)d,= -c,/[co(co + c,)]. Solution. Clearly, col is positive definite. Moreover, as a consequence of Lemma 14.2.17, Ai A* are nonnegative definite, and hence ciAi c^Ajt are
192 18. Sums (and Differences) of Matrices nonnegative definite. Thus, it follows from Corollary 14.2.5 that cq\ + £l=1 c/A,- is positive definite (and hence, in light of Lemma 14.2.8, nonsingular). That (cQl + £?=i c/A/)-1 = ^ol + H/=i 4? A,- is clear upon observing that (c0I + 2>A,)WbI + 2>A,) i i = c0dQI + co ^ diA,- + Jo 5^ c/A,- + ^ c/^/A? + ^ c/tf/A/A,- i i" / '.y's6' = I-y]-^-A/ + y^Al--V—^ A/+0 - T V^ C°C< ~ C<^C° + C'^ + C<^ A V Co(c0+cf-) = 1. EXERCISE 21. Let Ai,..., Ajt represent n x n symmetric idempotent matrices such that (for j ^ i = 1 k) A/Aj = 0, and let A represent an n x n symmetric idempotent matrix such that (for i = 1 A) C(A/) C C(A). Show that if rank(Ai) -\ 1- rank(Ajt) = rank(A), then AH 1- A* = A. Solution. Suppose that rank(Ai) -\ h rank(Ajt) = rank(A). Then, in light of Corollary 10.2.2, we have that tr(A, H +Ajt) = tr(Ai) + ...+ tr(A*) = rank(Aj) -\ h rank(Ajt) = rank(A) = tr(A). Moreover, [since C(A,-) C C(A)] there exists a matrix L/ such that A/ = AL/, so that AA,- = A2L/ = AL,- = A/ and A/A = AjA' = (AA/)' = Aj = A/ (i = 1,...,/:), implying that (A - £ A/)'(A - £ A,) = (A - £ A/)(A - £ A/) i i / i = A2 - J>A - £ AA/ + E A? + E A'A; i i i i.j^i = A-EAi-EA,+EA,+0 / i i = a-Ea„ i Thus, tr[(A - £A/)'(A - £ A/)] = tr(A - £ A/) = tr(A) - tr(]T A/) = 0. » i i i We conclude (on the basis of Lemma 5.3.1) that A — £. A/ = 0 or equivalently thatAi +...+A|t=A.
18. Sums (and Differences) of Matrices 193 EXERCISE 22. Let A represent an m x n matrix and B an n x m matrix. If B is a generalized inverse of A, then rank(I - BA) = « - rank(A). Show that the converse is also true; that is, show that if rank(I - BA) = n - rank(A), then B is a generalized inverse of A. Solution. Suppose that rank(I - BA) = n — rank(A). Then, since rank(BA) < rank(A), we have that n - rank(BA) > n - rank(A) = rank(I - BA). (S. 10) Moreover, making use of Corollary 4.5.9, we find that rank(I - BA) + rank(BA) > rank[(I - BA) + BA] = rank(I„) = n and hence that rankd - BA) > n - rank(BA). (S. 11) Together, results (S.10) and (S.l 1) imply that rank(I - BA) = n - rank(BA) or equivalently that rank(BA) + rank(I - BA) = n. Thus, it follows from Lemma 18.4.2 that BA is idempotent. Further, n - rank(A) = rank(I - BA) = n - rank(BA), implying that rank(BA) = rank(A). We conclude (on the basis of Theorem 10.2.7) that B is a generalized inverse of A. EXERCISE 23. Let A represent the (n x n) projection matrix for a subspace U ofK"xl along a subspace V of 11"xl (whereU0 V = 11"xl), and let B represent the (n x n) projection matrix for a subspace W of 1Z"xl along a subspace X of 7enxl(whereW0A' = 7enxl). (a) Show that A+B is the projection matrix for some subspace C of 11" x l along some subspace M ofll"*1 (where £@M = 1l"xl) if and only if BA = AB = 0, in which case C = U 0 W and M = V n X. (b) Show that A—B is the projection matrix for some subspace C of %" x l along some subspace M of ft"xl (where£0jM = ft" xl) if and only ifBA = AB = B, in which case C=UC\X and M = V@W. [Hint. Observe (in light of the result that a matrix is a projection matrix for one subspace along another if and only if it is idempotent and the result that a matrix, say K, is idempotent if and only if I — K is idempotent) that A — B is the projection matrix for some subspace C along some subspace M if and only if I — (A — B) = (I — A) + B is the projection matrix for some subspace C* along some subspace M*, and then make use of Part (a) and the result that the projection matrix for a subspace M* along a subspace C*
194 18. Sums (and Differences) of Matrices (where M* © C* = ft"x *) equals I - H, where H is the projection matrix for C* along M*. ] Solution, (a) According to Theorem 17.6.14, A and B are idempotent, U = C(A), W = C(B), V = MA), and X = Af(B). Suppose now that A + B is the projection matrix for some subspace C along some subspace M. Then, as a consequence of Theorem 17.6.13 (or 17.6.14), A+B is idempotent. And, it follows from Lemma 18.4.3 that BA = AB = 0. Conversely, suppose that BA = AB = 0. Then, as a consequence of Lemma 18.4.3, A+B is idempotent. And, it follows from Theorem 17.6.13 that A+B is the projection matrix for some subspace C along some subspace M and from Theorem 17.6.12 that C = C(A+B) and (in light of Theorem 11.7.1) that M =Af(A+B). Moreover, since A, B, and A + B are all idempotent, it follows from Theorem 18.4.1 that rank(A + B) = rank(A) + rank(B). (S.12) Since (according to Lemmas 4.5.8 and 4.5.7) rank(A + B) < rank(A, B) < rank(A) + rank(B), we have [as a consequence of result (S.12)] that rank(A + B) = rank(A, B) = rank(A) + rank(B), implying [in light of result (4.5.5)] that C(A + B) = C(A, B) and (in light of Theorem 17.2.4) that C(A) and C(B) are essentially disjoint. Thus, in light of result (17.1.4), it follows that C(A + B)=C(A)©C(B) or equivalently that C = U © VV. It remains to show that M = V C\ X or equivalently that N(A + B) = N(\) H MB). Let x represent an arbitrary vector in M(\ + B). Then, Ax + Bx = 0, and consequently (since A2 = A, B2 = B, and BA = AB = 0) Ax = A2x = A2x + ABx = A(Ax + Bx) = 0, Bx = B2x = B2x + BAx = B(Ax + Bx) = 0. Thus,x e Af(A)nAf(B). We conclude that^A+B) cAf(A)nM(B) and hence [since clearly M(\) n Af(B) C N(\ + B)] that ^(A + B) = N(\) n M(B). (b) Since (according to Lemma 10.1.2) I - (A - B) is idempotent if and only if A - B is idempotent, it follows from Theorem 17.6.13 that A - B is the projection matrix for some subspace C along some subspace M if and only if I — (A - B) = (I — A) + B is the projection matrix for some subspace C* along some subspace M*— Lemma 10.1.2 and Theorem 17.6.13 are the results mentioned parenthetically in the hint. And, since (according to Theorem 17.6.10) I - A is the
18. Sums (and Differences) of Matrices 195 projection matrix for V along U, it follows from Part (a) that (I - A) + B is the projection matrix for some subspace £* along some subspace M* if and only if B(I-A) = (I-A)B = 0, or equivalently if and only if BA = AB = B, in which case C* = V © W and M* = U OX. The proof is complete upon observing (in light of Theorem 17.6.10, which is the result whose use is prescribed in the hint) that if (I — A) + B is the projection matrix for V © W along WflA', then A - B = I - [(I - A) + B] is the projection matrix for U C\ X along V © W. EXERCISE 24. (a) Let B represent an n x n symmetric matrix, and let W represent an n x n symmetric nonnegative definite matrix. Show that WBWBW = WBW <& (BW)3 = (BW)2 & tr[(BW)2] = tr[(BW)3] = tr[(BW)4]. (b) Let Ai Ajt represent nxn matrices, let V represent an n x n symmetric nonnegative definite matrix, and define A = A] -\ A*. If VA/VA/V = VA/V for all / and if VA/VAyV = 0 for all / and ; # /, then VAVAV = VAV and rank(VAiV) + • • ■ + rank(VAjtV) = rank(VAV). Conversely, if VAVAV = VAV, then each of the following three conditions implies the other two: (1) VA/VAyV = 0 (for./^ / = 1 k) and rank(VA,VAl-V) = rank(VAIV) (for / = 1 *); (2) VA/VA/V = VA/V (for / = 1 k)\ (3) rank(VAiV) + h rank(VA*V) = rank(VAV). Indicate how, in the special case where Aj, ..., A* are symmetric, the conditions VAVAV = VAV and VA/VA/V = VA/V can be reexpressed by applying the results of Part (a) (of the current exercise). Solution, (a) Let S represent any matrix such that W = S'S — the existence of such a matrix follows from Corollary 14.3.8. If WBWBW = WBW, then clearly BWBWBW = BWBW, or equivalently (BW)3 = (BW)2. Conversely, suppose that (BW)3 = (BW)2. Then, (SB)'SBWBW = (SB)'SBW, implying (in light of Corollary 5.3.3) that SBWBW = SBW, so that WBWBW = S'SBWBW = S'SBW = WBW. It remains to show that (BW)3 = (BW)2 ^ tr[(BW)2] = tr[(BW)3] = tr[(BW)4]. Suppose that (BW)3 = (BW)2. Then, clearly, (BW)4 = BW(BW)3 = BW(BW)2 = (BW)3.
196 18. Sums (and Differences) of Matrices Thus, tr[(BW)2] = tr[(BW)3] = tr[(BW)4]. Conversely, suppose that tr[(BW)2] = tr[(BW)3] = tr[(BW)4]. Then, making use of Lemma 5.2.1, we find that tr[(SBS' - SBWBS')'(SBS' - SBWBS')] = tr(SBWBS' - 2SBWBWBS' + SBWBWBWBS') = tr(BWBS'S) - 2tr(BWBWBS'S) + tr(BWBWBWBS'S) = tr[(BW)2] - 2tr[(BW)3] + tr[(BW)4] = 0. Thus, it follows from Lemma 5.3.1 that SBS' - SBWBS' = 0, or equivalently that SBWBS' = SBS', so that (BW)3 = BS'(SBWBS')S = BS'(SBS')S = (BW)2. (b) Suppose that A] A* are symmetric (in which case A is also symmetric). Then, applying the results of Part (a) (with A in place of B and V in place of W), we find that VAVAV = VAV <S> (AV)3 = (AV)2 <S> tr[(AV)2] = tr[(AV)3] = tr[(AV)4]. Similarly, applying the results of Part (a) (with A,- in place of B and V in place of W), we find that VA/VA/V = VA,V <a> (A/V)3 = (A,V)2 & tr[(A,V)2] = tr[(A/V)3] = tr[(A/V)4]. EXERCISE 25. Let R represent an n x q matrix, S an n x m matrix, T an m x p matrix, and U a p x q matrix. (a) Show that rank(R + STU)=rank^ ~£TJ - rank(T). (E.3) (b)LetE/? = I-RR-.F/? = I-R"R,X = E/?ST,Y = TUF^Er = I-YY", Fx = I - X~X, Q = T + TUR-ST, and Z = EKQF*. Use the result of Part (b) of Exercise 10.10 to show that rank(R + STU) = rank(R) + rank(X) + rank(Y) + rank(Z) - rank(T). Solution, (a) Observing that /R -STV I 0\_/R + STU -ST\ \TV T )\-V l)~\ 0 T )
18. Sums (and Differences) of Matrices 197 and making use of Lemma 8.5.2 and Corollary 9.6.2, we find that rankQ^ ~£T)=rarik(R+0S™ ~£T) = rank(R + STU) + rank(T) and hence that rank(R + STU) = rankQ^ ~£T) - rank(T). Or, alternatively, equality (E.3) can be validated by making use of result (9.6.1) — observing that C(TU) C C(T) and ft(-ST) C ft(T), we find that rankL^ ~|T) = rank(T) + rank[R - (-ST)T-TU] = rank(T) + rank(R + STU) and hence that rank(R + STU) = rank^ ~jjf\ - rank(T). (b) Upon observing that rank(X) = rank(—X), that —X~ is a generalized inverse of -X, and that Fx = I - (-X-)(-X), it follows from the result of Part (b) of Exercise 10.10 that rankf J?T ~£A ) = rank(R) + rank(X) + rank(Y) + rank(Z). \TV T ) ~ We conclude, on the basis of Part (a) (of the current exercise) that rank(R + STU) = rank(R) + rank(X) + rank(Y) + rank(Z) - rank(T). EXERCISE 26. Show that, for any m x n matrices A and B, rank(A + B) > | rank(A) - rank(B) |. Solution. Making use of results (4.5.7) and (4.4.3), we find that rank(A) = rank[(A+B)-B] < rank(A+B)+rank(-B) = rank(A+B)+rank(B) and hence that rank(A + B) > rank(A) - rank(B). (S. 13) Similarly, we find that rank(B) = rank[(A+B)-A] < rank(A+B)+rank(-A) = rank(A+B)+rank(A)
198 18. Sums (and Differences) of Matrices and hence that rank(A + B) > rank(B) - rank(A) = -[rank(A) - rank(B)]. (S.14) Together, results (S.13) and (S.14) imply that rank(A + B) > |rank(A) - rank(B)|. EXERCISE 27. Show that, for any n x « symmetric nonnegative definite matrices A and B, C(A + B) = C(A, B), ft(A + B) rank(A + B) = rank(A -Kb)' ,B) = rank^Y Solution. According to Corollary 14.3.8, there exist matrices R and S such that A = R'R and B = S'S. And, upon observing that '♦■-GO© and recalling Corollaries 7.4.5 and 4.5.6, we find that C(A, B) = C(R'R, S'S) = C(R', S') = dY*) 1 = C(A + B) [which implies that rank(A, B) = rank(A + B)] and similarly that Kb)=Ks;s)=Ks)=K(A+B) [which implies that rank! R I = rank(A + B)]. EXERCISE 28. Let A and B represent m x n matrices. (a) Show that (1) C(A) C C(A + B) if and only if rank(A, B) = rank(A + B) and (2) ft(A) c K{\ + B) if and only if rankf R) = rank(A + B). (b) Show that (1) if 71(A) and 71(B) are essentially disjoint, then C(A) C C(A+ B) and (2) if C(A) and C(B) are essentially disjoint, then 7£(A) c 11(A + B). Solution, (a)(1) Suppose that rank(A, B) = rank(A+B). Then, since (according to Lemma 4.5.8) C(A+B) CC(A,B), it follows from Theorem 4.4.6 that C(A,B) = C(A + B). Then, since C(A) C C(A, B), we have that C(A) C C(A + B). Conversely, suppose that C(A) C C(A + B). Then, according to Lemma 4.2.2, there exists a matrix F such that A = (A + B)F. Further, B = (A + B) - A =
18. Sums (and Differences) of Matrices 199 (A + B)(I - F). Thus (A, B) = (A + B)(F, I - F), implying that C(A, B) C C(A + B). Since (according to Lemma 4.5.8) C(A + B) C C(A, B), we conclude that C(A, B) = C(A + B) and hence that rank(A, B) = rank(A + B). (2) The proof of Part (2) is analogous to that of Part (1). (b) Let c = dim[C(A) n C(B)], d = dim[ft(A) fl ft(B)], and «=[■-(;)©"](: ii^-^nu,]. (1) If 1Z(A) and 1Z(B) are essentially disjoint (or equivalently if d = 0), then (as a consequence of Theorem 18.5.6) rank(H) = 0 = dy implying [in light of result (5.14)] that rank(A + B) = rank(A, B), and hence [in light of part (a)-( 1)] that C(A) C C(A + B). (2) Similarly, if C(A and C(B) are essentially disjoint (or equivalently if c = 0), then (as a consequence of Theorem 18.5.6) rank(H) = 0 = c, implying [in light of result (5.18)] that rank(A + B) = rank(gj and hence [in light of Part (a)-(2)] that 11(A) C ft(A + B). EXERCISE 29. Let A and B represent //; x n matrices. Show that each of the following five conditions is necessary and sufficient for rank additivity [i.e., for rank( A + B) = rank(A) + rank(B)]: (a) rank(A, B) = rank( £ J = rank(A) + rank(B); (b) rank(A) = rank[A(I - B~B)] = rank[(I - BB~)A]; (c) rank(B) = rank[B(I - A"A)] = rank[(I - AA~)B]; (d) rank(A) = rank[A(I - B~B)] and rank(B) = rank[(I - AA~)B]; (e) rank(A) = rank[(I - BB~)A] and rank(B) = rank[B(I - A-A)]. Solution. Let r = dim[C(A) fl C(B)] and s = dim[1Z(\) fl 11(B)]. In light of Theorem 18.5.7, it suffices to show that the condition r = s = 0 is equivalent to each of Conditions (a)-(e). That r = s = 0 implies Condition (a) and conversely is an immediate consequence of results (5.15) and (5.19). That r = s = 0 is equivalent to each of Conditions (b) - (e) becomes clear upon observing that (as a consequence of Corollary 17.2.10) r = 0 <& rank(A) = rank[(I-BB~)A] <S> rank(B) = rank[(I - AA~ )B],
200 18. Suras (and Differences) of Matrices s=0 & rank(A) = rank[A(I-B~B)] <& rank(B) = rank[B(I - A-A)]. EXERCISE 30. Let A and B represent mxn matrices. And, let -['-(SX$)1(i 1)*-™-™ (a) Show that rank(A - B) = rank(A) - rank(B) + [rank(A, B) - rank(A)] + [rank( ^ J - rank(A)] + rank(H). Do so by applying (with —B in place of B) the formula rank(A + B) = rank(A, B) + rank(£ J - rank(A) - rank(B) + rank(K), (*) where *=['-(«)(«)"](£ ly™-™ (b) Show that A and B are rank subtractive [in the sense that rank(A — B) = rank(A) — rank(B)] if and only if rank(A, B) = rankl R I = rank( A) and H = 0. (c) Show that if rank(A.B) = rank(B J = rank(A), then (1) (A",0) and I ft ) are generalized inverses of ( R J and (A, B), respectively, and (2) for (y" = (A-,0)and(A.B)- = (f) H=(j BA_»_B). (d) Show that each of the following three conditions is necessary and sufficient for rank subtractivity [i.e., for rank(A — B) = rank(A) — rank(B)]: (1) rank(A,B) = rank(gj=rank(A) and BA~B = B; (2) C(B) c C(A), 11(B) c ft(A), and BA~B = B; (3) AA-B = BA~A = BA-B = B. (e) Using the result of Exercise 29 (or otherwise), show that rank(A - B) = rank(A) - rank(B) if and only if rank(A - B) = rank[A(I - B~B)] = rank[(I - BB")A]. Solution, (a) Clearly,
18. Sums (and Differences) of Matrices 201 Thus, it follows from Parts (1) and (2) of Lemma 9.2.4 that ( £ J H" _J \ is a generalized inverse of I ) and that ( " _- j (A, B)~ is a generalized inverse of (A, —B). Further, ¢- in -i) -[-(4)6)-(5- -1)1 *(S -i)["-fr _',.)"'«..>-«.-.,} Now, applying result (*) [which is equivalent to result (5.7) or, when combined with result (17.4.13) or (17.4.12), to result (5.8) or (5.9)] and recalling Corollary 4.5.6, we find that rank(A - B) = rank[A + (-B)] = rank(A, -B)+rank( _) - rank(A) - rank(-B) -[fr-lK-l)] ¢) ink(i 0- = rank(A, B) + rank! R 1 - rank(A) - rank(B) + rank(H) = rank(A) - rank(B) + [rank(A, B) - rank(A)] + [rank( B J - rank(A)] + rank(H). (b) It follows from Part (a) that rank(A - B) = rank(A) - rank(B) if and only if [rank(A, B) - rank(A)] + [rankf £ J - rank(A)] + rank(H) = 0. (S.15) Since all three terms of the left side of equality (S.15) are nonnegative, we conclude that rank(A — B) = rank(A) — rank(B) if and only if rank(A, B) — rank(A) = 0, rank(R ) - rank(A) = 0, and rank(H) = 0 or, equivalently, if and only if rank(A, B) = rankf £ j = rank(A) and H = 0.
202 18. Sums (and Differences) of Matrices (c) Suppose that rank(A, B) = rank! R) = rank(A). Then, according to Corollary 4.5.2, C(B) C C(A) and 11(B) C ft(A), and it follows from Lemma 9.3.5 that AA~B = B and BA~A = B. Thus,(l) eK--e)-G)™-(£*)-(i> (A, B)(£ Va, B) = AA~(A, B) = (AA~A, AA"B) = (A, B); and (2) upon setting I R ) and (A, B)~ equal to (A~, 0) and I ft I, respectively, we obtain -[•-CMC ■«>»] _/I-AA~ 0\/A 0\/I-A~A -A~B\ ~^-BA~ l)\0 -h)\ 0 I ) /0 0\/I-A~A -A"B\ ~ ^-BA~A -BJ^ 0 I ) \0 BA-AA-B-B/ ^0 BA~B-B/ (d)( )(1) Suppose that rank(A, B) = rank ( B) = rank(A) and that BA~B = B. Then, it follows from Part (c) that (A~, 0) and I ft J are generalized inverses of (R ) and (A, B), respectively, and that, for ( J = (A~, 0) and (A, B)~ = 10 ), H = 0. Thus, as a consequence of Part (b), we have that rank( A - B) = rank(A) - rank(B). Conversely, suppose that rank(A — B) = rank(A) — rank(B). Then, according to Part (b), rank(A, B) = rank I R ) = rank(A) and H = 0 [for any choice of I R J and (A, B)"]. And, observing [in light of Part (c)J that (A~, 0) and I 0 J are generalized inverses of I R j and (A, B), respectively, and that, for (J) = (A_- °> «* (A'B)~ = (o")• H = (J BAB - b)' We find that BA" B - B = 0 or equivalently that BA"B = B.
18. Sums (and Differences) of Matrices 203 (2) Since (according to Corollary 4.5.2) C(B) C C(A) & rank(A, B) = rank(A)andft(B) C 11(A) <& rankf ) = rank(A), Condition (2) is equivalent to Condition (1) and hence is necessary and sufficient for rank subtractivity. (3) Since (according to Lemma 9.3.5) AA~B = B <s> C(B) C C(A) and BA~A = B <s> 7?(B) C 11(A). Condition (3) is equivalent to Condition (2) and hence is necessary and sufficient for rank subtractivity. (e) Clearly, rank( A-B) = rank(A) - rank(B) if and only if rank[(A-B) + B] = rank(A - B) + rank(B), that is, if and only if A - B and B are rank additive. Moreover, it follows from the result of Exercise 29 [specifically. Condition (b)] that A — B and B are rank additive if and only if rank(A - B) = rank[(A - B)(I - B"B)] = rank[(I - BB")(A - B)]. Since(A-B)(I-B-B) = A(I-B-B)and(I-BB-)(A-B) = (I-BB-)A, we conclude that rank(A — B) = rank( A) — rank(B) if and only if rank(A — B) = rank[A(I - B~B)] = rank[(I - BB~)A]. EXERCISE 31. Let A i A* represent m x n matrices. Adopting the terminology of Exercise 17.6, use Part (a) of that exercise to show that if 7£(Aj), ..., 72.(Ajt) are independent and C(A\) C(A*) are independent, then rank(Aj + 1- A*) = rank(Aj H 1- rank(Ajt). Solution. Suppose that TZ(A\) ft(Ajt) are independent and also that C(A\), ..., C(Ajt) are independent. Then, as a consequence of Part (a) of Exercise 17.6, we have that, for / = 2 A\ ft(A/) and ft(Aj ) + ••• + ft(A/_i) are essentially disjoint and C(A,) and C(Ai) -\ 1- C(A,_i) are essentially disjoint or equivalent^ [in light of results (17.1.7) and (17.1.6)] that, for / = 2,..., K ft(A,-) and /A, \ H\'- I are essentially disjoint and C(A,-) andC(Ai A/_j) are essentially U-i/ disjoint. Moreover, it follows from results (4.5.9) and (4.5.8) that (for/ = 2 k) /A, X K(Ai+-- + A,_i)c72 : I and C(Ai+--+A,_i) cC(A| A,_i). U-i/ Thus, for i = 2 k, ft(A,) and 1Z(A\ -\ + A,_i) are essentially disjoint and C(A,) and C(A\ -\ h A,_i) are essentially disjoint, implying (in light of Theorem 18.5.7) that (for / = 2 k) rank(Ai + • • • + A/_i + A,-) = rank(Aj + • • • + A/_i) + rank(A/). We conclude that rank(Ai + • • • + A*) = rank(Aj + • • • + Ajt_j) + rank(Ajt)
204 18. Sums (and Differences) of Matrices = rank(Ai + • • • + Ajt_2) + rank(Ajt-i) + rank(Ajt) = rank(Ai) -\ h rank(Ajt). EXERCISE 32. Let T represent an m x p matrix, U an m x q matrix, V an n x p matrix, and W an n x q matrix, and define Q = W — VT~U. Further, let Er=I-TT-,Fr = I-T-T, X = ErU,andY = VFr. (a) Show that rankfj ^ = rank(T) + rank^ *\ (E.4) (b) Show that rankfu ^ = rank(T) + rank^"^ U *\ [Hint. Observe that (since the rank of a matrix is not affected by a permutation of (0 V\ /T U\ ,, T I = rankf v ft J, and make use of Part (a).] (c) Show that rankf f°T Z ) = rank(T) + rank(X) + rank(Y) + rank(EyVT~UFx), where Ey = I - YY~ and F* = I - X~X. Hint. Use Part (b) in combination with the result that, for any r x s matrix A,rxu matrix B, and it x s matrix C, <c 9- rank )= rank(B) + rank(C) + rank[(I - BB")A(I - C~C)]. (*) (d) Show that c(v w)=I rank(v w)=rank(T) + rank(Q) + rank(A) + rank(B) + rank[(I-AA-)XQ-Y(I-B-B)], where A = X(I- Q~Q) and B = (I- QQ~) Y. [Hint. Use Part (c) in combination with Part (a).] Solution, (a) Let K = Q - YT"U. Then, clearly, ( I 0\ /T U\ /I -T~U\ _ /T U\ /I -T-1A ^-vt- i)\y w)\0 I j-^Y QJl^O I ) -e 9-
18. Sums (and Differences) of Matrices 205 so that (in light of Lemma 8.5.2) rank(v w)=rank(Y K/' (S*16) Further, /T X\_/T 0\ /0 X\ [y Ky~^o 07 +\Y k)' And, since 111 ft J = 7£(T) and Til Y J = 7£(Y) and since (according to Corollary 17.2.8) ft(T) and ft(Y) are essentially disjoint, Til 0 J and Tily) are es- /T 0\ sentially disjoint, and hence it follows from Corollary 17.2.16 that ft I ft ft J (0 X\ /T 0\ /0 X\ Y K I are essentially disjoint. Similarly, Ci ft ft I and C( Y K I are essentially disjoint. Thus, making use of Theorem 18.5.7, we find that rankp *) = rank(J J)+ranlc(2 k) = rank(T) + rank^Y £)• (S'17) Now, observe that /0 X\/I T-U\_/0 X\ \Y *)\0 I )- \Y Q) rank(J §=«■*($ q)- (S.18) Finally, combining results (S.16HS.18), we obtain rank(Y ^=rank(T)+rank^ ^ = rank(T) + rank ^ q)- (b) Since (as indicated by Lemma 8.5.1) the rank of a matrix is not affected by a permutation of rows or columns, rank(S T)=rank(v J)' (S19) Moreover, applying result (E.4) (in the special case where W = 0) and again making use of Lemma 8.5.1, we find that rank(v o)=rankcr)+rank(Y -vr-u) = rank(T)+rarj/~V£ U *\ (S.20) and hence that
206 18. Sums (and Differences) of Matrices And, upon combining result (S.20) with result (S.19), we obtain /0 V\ /-VT~U Y\ rankljj TJ =rank(T) + rankl x 0J. (c) Applying result (*) [or equivalently result (17.2.15), which is part of Theorem 17.2.17] with -VT~U, Y, and X in place of A, B, and C, respectively (or T, U, and V, respectively), we find that rank( ^ U ^) = rank(Y) + rank(X) +rank[(I - YY-)(-VT~U)(I - X~X)] = rank(Y) + rank(X) + rank(Ey VT-UFx). (S.21) And, upon combining result (S.21) with Part (b), we obtain rank( u ^J = rank(T) + rank(Y) + rank(X) + rank(EyVT"UFx). (d) Applying Part (c) (with Q, Y, and X in place of T, U, and V, respectively, and hence with A and B in place of Y and X, respectively), we find that rank( Y * J = rank(Q) + rank(B) + rank(A) + rank[(I - AA~)XQ~ Y(I - B~B)]. (S.22) And, upon combining result (S.22) with Part (a), we obtain rank( y J = rank(T) + rank(Q) + rank(A) + rank(B) + rank[(I - AA~)XQ"Y(T - B~B)]. EXERCISE 33. Let R represent an n x q matrix, S an n x m matrix, T an m x p matrix, and U a p x q matrix, and define Q = T + TUR~ST. Further, let E* = I - RR-, F* = I - R-R, X = E*ST. Y = TUF*, A = X(I - Q~Q), B = (I — QQ~)Y. Use the result of Part (d) of Exercise 32 in combination with the result of Part (a) of Exercise 25 to show that rank(R + STU) = rank(R) + rank(Q) - rank(T) + rank(A) + rank(B) + rank[(I - AA~)XQ-Y(I - B"B)]. Solution. Upon observing that rank(A) = rank(-A), that -A" is a generalized inverse of —A, and that rank[(I - AA-)XQ~Y(I - B~B)] = rank[-(I - AA-)XQ~Y(I - B~B)] = rank{[I - (-A)(-A")](-X)Q-Y(I - B~B)},
18. Sums (and Differences) of Matrices 207 it follows from the result of Part (d) of Exercise 32 that (R —ST\ TO T J = rank(R) + rank(Q) + rank(A) + rank(B) + rank[(I - AA)XQY(I - B~B)]. We conclude, on the basis of Part (a) of Exercise 25, that rank(R + STU) = rank(R) + rank(Q) - rank(T) + rank(A) + rank(B) + rank[(I-AA~)XQ-Y(I-B-B)].
19 Minimization of a Second-Degree Polynomial (in n Variables) Subject to Linear Constraints EXERCISE 1. Let a represent annxl vector of (unconstrained) variables, and define /(a) = a'Va — 2b'a, where V is an n x n matrix and bannxl vector. Show that if V is not nonnegative definite or if b & C(V), then /(a) is unbounded from below, that is, corresponding to any scalar c, there exists a vector a* such that /(a*) < c. Solution. Let c represent an arbitrary scalar. Suppose that V is not nonnegative definite. Then, there exists annxl vector x such that x'Vx < 0. Moreover, for any scalar k, f(kx) = k2(x'\x) - 2k(b'x) = k[k(x'\x) - 2(b'x)]. Thus, limjt_^±oo f(kx) = -co, and hence there exists a scalar /:* such that f(k*x) < c, so that, for a* = /:+x, /(a*) < c. Or, suppose that b ¢ C(V). According to Theorem 12.5.11, C(V) contains a (unique) vector bj and C^CV) [or equivalently Af(V')] contains a (unique) vector b2 such that b = b\ + b2. Clearly, b2 £ 0 [since otherwise we would arrive at a contradiction of the supposition that b ¢ C(V)]. Moreover, for any scalar k, f(kb2) = ^2(V,b2),b2 - 2*b'b2 = -2k(b\ + b2)b2 = -2/r(b2b2). Thus, limjt-^oo f(kb2) = -co, and hence there exists a scalar /:* such that f(k*b2) < c, so that, for a* = **b2, /(a*) < c. EXERCISE 2. Let V represent an n x n symmetric matrix and X an n x p matrix. Show that, for any p x p matrix U such that C(X) C C(V + XUX7),
210 19. Minimization of a Second-Degree Polynomial (1) (V + XUX')(V + XUX')-V = V; (2) vcv+xux'r (v+xuxo = v. Solution. (1) According to Lemma 19.3.4, C(V, X) = C(V+XUX'), implying (in light of Lemma 4.5.1) that C(V) C C(V + XUX') and hence (in light of Lemma 9.3.5) that (V + XUX')(V + XUX'rV = V. (2) According to Lemma 19.3.4, C(V, X) = C(V + XU'X'), implying that C(V) C C(V + XU'X') and hence (in light of Lemma 4.2.5) that ft(V) C ft[(V + XU'X')'] = ft(V + XUX'). Thus, it follows from Lemma 9.3.5 that v(v+xux'r (v+xux') = v. EXERCISE 3. Let V represent an n x /z symmetric nonnegative definite matrix, Xannxp matrix, B an n x s matrix such that C(B) c C(V, X), and Da px s matrix such that C(D) C C(X'). Further, let U represent any p x p matrix such that C(X) C C(V + XUX'), and let W represent an arbitrary generalized inverse of V + XUX'. Devise a short proof of the result that A* and R* are respectively the first (« x s) and second (p x s) parts of a solution to the (consistent) linear system ff »)©=© (in an n x s matrix A and a p x s matrix R) if and only if R* = T* + UD and A* = WB - WXT* + [I - W(V + XUX')]L for some solution T* to the (consistent) linear system X'WXT = X'WB-D (**) (in a p x s matrix T) and for some n x s matrix L. Do so by taking advantage of the result that if the coefficient matrix, right side, and matrix of unknowns in a linear system HY = S (in Y) are partitioned (conformally) as -OS S> S=(S> - *-©■ and if C(H12)CC(Hn), C(Si)cC(Hn), and ft(H2l) C ft(Hn),
19. Minimization of a Second-Degree Polynomial 211 then the matrix Y* = ( J j is a solution to the linear system HY = S if and only if Y, is a solution to the linear system (H22 - H21H-Hi2)Y2 = S2 - H21H-Sj (in Y2) and YJ and Y, are a solution to the linear system Hi 1 Yi + H12Y2 = Si (in Yi and Y2>. Solution. According to Lemma 19.3.2, A* and R* are the first and second parts of a solution to linear system (*) [or equivalently linear system (3.14)] if and only if A* and R* — UD are the first and second parts of a solution to the linear system CT* S)(i)-© (in A and T). Now, observing (in light of Lemma 19.3.4) that ft(X') C ft(V + XUX') and C(B) C C(\ + XUX'), it follows from the cited result [or equivalently from Part (1) of Theorem 11.11.1 ] that A* and T* are the first and second parts of a solution to linear system (S.l) if and only if T* is a solution to the linear system (0 - X'WX)T = D - X'WB (S.2) (in T) and (V + XUX')A*+XT* = B. (S.3) Note that linear system (S.2) is equivalent to linear system (**) [which is identical to linear system (3.17)]. Note also that condition (S.3) is equivalent to the condition (V + XUX')A* = B - XT*. (S.4) And, since (in light of Lemma 19.3.4) C(B - XT*) c C(V + XUX'), it follows from Theorem 11.2.4 that condition (S.4) is equivalent to the condition that A* = WB - WXT* + [I - W(V + XUX')]L for some matrix L. Thus, A* and R* — UD are the first and second parts of a solution to linear system (S.l) if and only if R* - UD = T* (or equivalently R* = T* + UD) and A* = WB - WXT* + [I - W(V + XUX')]L for some solution T* to linear system (**) [or equivalently linear system (3.17)] and for some matrix L. EXERCISE 4. Let V represent an n x n symmetric matrix and X an n x p matrix. Further, let U = X'TT'X, where T is any matrix whose columns span the null
212 19. Minimization of a Second-Degree Polynomial space of V. Show that C(X) C C(V + XUX') and that C(V) and C(XUX') are essentially disjoint and 7£(V) and ft(XUX') are essentially disjoint (even in the absence of any assumption that V is nonnegative definite). Solution. Making use of Corollaries 7.4.5,4.5.6, and 17.2.14, we find that C(V, X) = C(V, XX') = C(V, XX'T). Further, C(XX'T) = C[(XX'T)(XX'T)'] = C(XUX'). Thus, again making use of Corollary 4.5.6, we have that rank(V, X) = rank(V, XX'T) = rank(V, XUX'). And, in light of Corollary 17.2.14, C(V) and C(XUX') are essentially disjoint. Moreover, since XUX' is clearly symmetric, it follows from Lemma 17.2.1 that TZ(\) and ft(XUX') are also essentially disjoint. Finally, in light of Theorem 18.5.6, it follows from result (5.13) or (5.14) that rank(V + XUX') = rank(V, XUX'), so that rank(V + XUX') = rank(V, X) or equivalently (in light of Lemma 19.3.4) C(X) C C(V + XUX'). EXERCISE 5. Let Vrepresent annxn symmetric nonnegative definite matrix and X an n x p matrix. Further, let Z represent any matrix whose columns span J\f(X') or, equivalently, C±(X). And, adopt the same terminology as in Exercise 17.20. (a) Using the result of Part (a) of Exercise 17.20 (or otherwise) show that an n x n matrix H is a projection matrix for C(X) along C(VZ) if and only if H' is the first (/2 x n) part of a solution to the consistent linear system (in an n x n matrix A and a p x n matrix R). (b) Letting U represent any p x p matrix such that C(X) C C(V + XUX;) and letting W represent an arbitrary generalized inverse of V + XUX\ show that an n x n matrix H is a projection matrix for C(X) along C(VZ) if and only if H = Px.w + K[I - (V + XUX')W] for some n x n matrix K. Solution, (a) In light of the result of Part (a) of Exercise 17.20, it suffices to show that HX = X and HVZ = 0 (or equivalently that X'H' = X' and Z'VH' = 0) if and only if H; is the first part of a solution to linear system (E.l). Now, suppose that X'H' = X; and Z'VH' = 0. Then, in light of Corollary 12.1.2, it follows from Corollary 12.5.5 that C(VH/) c C(X) and hence (in light
19. Minimization of a Second-Degree Polynomial 213 of Lemma 4.2.2) that VH' = XT for some matrix T. Thus, (x' o)(-t) = (x')' so that H' is the first part of a solution to linear system (E.1). Conversely, suppose that H' is the first part of a solution to linear system (E.1) and hence that (x' o)(rJ = (x') for some pxn matrix R*. Then, X'H' = X'. And, VH' = X(-R„), implying that C(VH') C C(X) and hence (in light of Corollaries 12.5.5 and 12.1.2) that Z'VH' = 0. (b) In light of Part (a), it suffices to show that H' is the first part of a solution to linear system (E. 1) if and only if H = Px.w + K[I - (V + XUX')W] for some matrix K. According to Lemma 19.3.4, C(X) C C(\ + XU'X'). Moreover, since V + XU'X' = (V+XUX')' and since X'W'X = (X'WX)', W is a generalized inverse of V+XU'X\ and [(X'WX)"]' is a generalized inverse of X'W'X. Thus, it follows from the results of Section 19.3c that H' is the first part of a solution to linear system (Rl)ifandonlyif H' = W,X[(X,WX)-],X/ + [I - W(V + XU'X')]K' for some (nxn) matrix K, or equivalently if and only if H = Px.w + K[I - (V + XUX') W] for some matrix K. EXERCISE 6. Let V represent an n x n symmetric nonnegative definite matrix, W an n x n matrix, and X an n x p matrix. Show that, for the matrix WX(X/WX)~X/ to be the first (nxn) part of some solution to the (consistent) linear system (j J) ¢)-(¾ (in an n x n matrix A and a p x n matrix R), it is necessary and sufficient that C(VWX) C C(X) and rank(X,WX) = rank(X). Solution. Clearly, WX(X/WX)~X/ is the first part of a solution to linear system (E.2) if and only if VWX(X/WX)~X/ + XR = 0 for some matrix R and X,WX(X,WX)-X/ = X\ or equivalently if and only if ctvwxtx'wxrx'] c C(X) (S.5)
214 19. Minimization of a Second-Degree Polynomial and X'WX(X'WX)-X' = X'. (S.6) And, by following the same line of reasoning as in the latter part of the proof of Theorem 19.5.1, we find that conditions (S.5) and (S.6) are equivalent to the conditions that C(VWX) c C(X) and rank(X'WX) = rank(X). EXERCISE 7. Let a represent an n x 1 vector of variables, and impose on a the constraint X;a = d, where X is an n x p matrix and d is a p x 1 vector such that d e C(X'). Define /(a) = a'Va — 2b'a, where V is an n x n symmetric nonnegative definite matrix and b is an n x 1 vector such that b 6 C(V, X). Further, define g(a) = a'(V + W)a - 2(b + c)'a, where W is any n x n matrix such that C(W) C C(X) and TZ(W) C 1Z(X!) and where c is any n x 1 vector in C(X). Show that the constrained (by X'a = d) minimization of g(a) is equivalent to the constrained minimization of /(a) [in the sense that g(a) and /(a) attain their minimum values at the same points]. Solution. Clearly, c = Xr for some p x 1 vector r. Further, in light of Lemma 9.3.5, we have that W = XX" W and W = W(X')~X' and hence that W = XX-W(X,)~X/ = XUX', where U = X-W(X')~. Thus, g(a) = /(a)+3'W3-2c'3 = /(3) + (X'3)'UX'3 - 21^3, so that, for 3 such that X'a = d, g(3) = /(3)+d,Ud-2r,d. We conclude that, for 3 such that X's = d, g(a) differs from /(3) only by an additive constant and hence that g(a) and /(a) attain their minimum values (under the constraint X'a = d) at the same points. EXERCISE 8. Let V represent an n x n symmetric nonnegative definite matrix, W an n x n matrix, X an n x p matrix, f an n x 1 vector, and d a p x 1 vector. Further, let b represent an n x 1 vector such that b 6 C(V, X). Show that, for the vector W(I - Px.w)f + WX(X'WX)-d to be a solution, for every d e C(X'), to the problem of minimizing the second-degree polynomial a'Va — 2b'a (in a) subject to X'a = d, it is necessary and sufficient that VWf-beC(X), (E.3) C(VWX) C C(X), (E.4) and rank(X'WX) = rank(X). (E.5)
19. Minimization of a Second-Degree Polynomial 215 Solution. It follows from Theorem 19.2.1 that a'Va — 2b'a has a minimum at W(I - Px.w)f+WX(X'WX)~d under the constraint X'a = d [where d e C(X')] ifandonly ifVW(I-Px.w)f+VWX(X,WX)-d+Xr = b for some vector r and X,W(I-Px.w)f+X,WX(X,WX)-d = d,orequivalently if and only ifVW(I- Px.w)f-b+VWX(X'WX)-d 6 C(X)andX'W(I-Px.w)f+X'WX(X'WXrd = d.Thust for W(I-Px.w)f+WX(X;WX)-dtobeasolution, for every d e C(X'), to the problem of minimizing a'Va - 2b'a subject to X'a = d, it is necesary and sufficientthat,forevery/zx 1 vectoru,VW(I-Px.w)f-b+VWX(X'WXrX'u e C(X)andX,W(I-Px.w)f+X,WX(X,WX)-X,u = X'u, a requirement equivalent to a requirement that VW(I-Px.w)f-beC(X), (S.7) ctvwxtx'wxrx'] C C(x>, cs.8) and X'WX(X'WX)-X' = X', (S.9) as we now show. Suppose that conditions (S.7)-(S.9) are satisfied. Then, observing that x'wpx.w = x'wxcx'wxrx'w, we find that, for every u, VW(I - Px,w)f - b + VWX(X'WXrX'u 6 C(X) and X'W(I - Px,w)f + X'WX(X'WX)-X'u = (X'W - X'W)f + X'u = X'u. Conversely, suppose that, for every u, VW(I - Px.w)f - b + VWX(X,WX)~X,u e C(X) (S.10) and x'wa - px,W)f+x'wxcx'wxrx'u = x'u. (s.in Then, since conditions (S.10) and (S.l 1) are satisfied in particular for u = 0, we have that VW(I-Px,w)f-beC(X) and X,W(I-Px.w)f=0. Further, for every u, VWX(X'WXrX'u e C(X) and X'WX(X'WX)-X'u = X'u, implying that crvwxcx'wxrx'i c ax) and X,WX(X,WX)-X/ = X'.
216 19. Minimization of a Second-Degree Polynomial Now, when condition (S.8) is satisfied, condition (E.3) is equivalent to condition (S.7), as is evident from Lemma 4.1.2 upon observing that VWPx,w = VWX(X'WX)-X'W and hence that VWPx.wf 6 C[VWX(X'WX)-X']. Moreover, by employing the same line of reasoning as in the latter part of the proof of Theorem 19.5.1, we find that conditions (E.4) and (E.5) are equivalent to conditions (S.8) and (S.9). Thus, conditions (E.3)-(E.5) are equivalent to conditions (S.7) - (S.9). And, we conclude that, for W(I - Px.w)b + WX(X'WX)~d to be a solution, for every d e C(X'), to the problem of minimizing a'Va - 2b'a subject to X'a = d, it is necessary and sufficient that conditions (E.3)-(E.5) be satisfied. EXERCISE 9. Let V and W represent n x n matrices, and let X represent an /2 x p matrix. Show that if V and W are nonsingular, then the condition C( VWX) C C(X) is equivalent to the condition C(V_,X) C C(WX) and is also equivalent to the condition ^(W-'V-'X) C C(X). Solution. Assume that V and W are nonsingular. Then, in light of Corollary 8.3.3, rank(VWX) = rank(X), rank^-'X) = rank(X) = rank(WX), and rank(W-1 V_1X) = rank(X). Thus, as a consequence of Theorem 4.4.6, C(VWX)CC(X) <* C(VWX)=C(X), C(V~lX) C C(WX) & C(\-lX)=C(WX)1 and CCW'v'X) c C(X) & C(W~1V-1X)=C(X). Now, if C(VWX)cC(X), then C(X) = C(VWX), so that X = VWXQ for some matrix Q, in which case V_1X = WXQ and W"1 V_1X = XQ, implying thatCKV-'X) c C(WX) andCKW-'V-'X) C C(X). Conversely, if C(\~lX) c C(WX), then C(WX) = C(\~lX), so that WX = V'XQ for some matrix Q, in which case VWX = XQ, implying that C(VWX) C C(X). And, similarly, if C(W-i v-lX) c C(X), then C(X) = C(W_1 V-'X), so that X = W"1 V^XQ for some matrix Q, in which case VWX = XQ, implying that C(VWX) C C(X). We conclude that C(VWX) C C(X) & C(\~lX) C C(WX) and that C(VWX) c C(X)<&C(W-1\-1X)CC(K). EXERCISE 10. Let V represent an n x n symmetric positive definite matrix, W an n xn matrix, X an n x p matrix, and d a p x 1 vector. Show that, for the vector WX(X;WX)~d to be a solution, for every d e C(X'), to the problem of minimizing the quadratic form a'Va (in a) subject to X'a = d, it is necessary and sufficient that V_1Px.w' be symmetric and rank(X'WX) = rank(X). Show that it is also necessary and sufficient that (I - Px w, )V~' Px.\v = 0 and rank(X'WX) = rank(X). Solution. Since (V'Px.w')' = Px.w'V-1' v~,px.w' is symmetric if and only if V_1Px.w = Px.w'V_1- Moreover, if V'Px.^ = P^V"1, then Px.w'V = V(V-,Px.w')V = V(PX^V-'JV = VPX w,.
19. Minimization of a Second-Degree Polynomial 217 And, conversely, if Px.w'V = VPXW,, then V-'Px.w = V-'CPx^V-1 = V-UVP'^V-1 =P^WV"1. Thus, V~lPx.w' is symmetric if and only if Px.w'V = VPX w,, and the necessity and sufficiency of V-1PX>W' being symmetric and rank(X'WX) = rank(X) follows from Theorem 19.5.4. To complete the proof, it suffices to show that (I — Pxw/)V_1Px.w' = 0 and rank(X'WX) = rank(X), or equivalently that V",Px.wf = px,w'V~lpx.W and rank(X'WX) = rank(X), if and only if V_1PX,W/ = PXW,V-1 and rank(X'WX) = rank(X). Suppose that V-1PXtW' = PX%WV_I and rank(X'WX) = rank(X). Then, since X'W'X = (X'WX)', rank(X'W'X) = rank(X), and it follows from Part (3) of Lemma 19.5.5 that Px w, = Pxw>. Thus, V-'P^ = (V-^xwOPxw = P^V-'P^. Conversely, if V-1PX>W' = PXW/V-1PX%W', then (since clearly the matrix Pxw,V_1Px.w' is symmetric) V-!PW = (P^V-^/ = (V-'Px.mt)' = P'x.w'V"1- We conclude that V_1Px,w' = px,wv~Ipx.W and rank(X'WX) = rank(X) if and only if \~lFx,w = Px.w'V_1 and rank(X'WX) = rank(X). EXERCISE 11. Let V represent an n x n symmetric nonnegative definite matrix, X an n x p matrix, and d a p x 1 vector. Show that each of the following six conditions is necessary and sufficient for the vector X(X'X)~d to be a solution, for every d 6 C(X'), to the problem of minimizing the quadratic form a'Va (in a) subject to X'a = d: (a) C(VX) C C(X) (or, equivalently, VX = XQ for some matrix Q); (b) PXV(I - Px) = 0 (or, equivalently, PXV = PXVPX); (c) PxV = VPX (or, equivalently, PXV is symmetric); (d) C(VPX) C C(PX); (e) C(VPx)=C(V)flC(Px); (f) C(VX)=C(V)flC(X). Solution, (a), (b), and(c) Upon applying Theorems 19.5.1 and 19.5.4 (with W = I) and recalling (from Corollary 7.4.5) that rank(X'X) = rank(X) and (from Theorem 12.3.4) that Px is symmetric, we find that each of Conditions (a)-(c) is necessary and sufficient for X(X'X)~d to be a solution to the problem of minimizing a'Va subject to X'a = d.
218 19. Minimization of a Second-Degree Polynomial (d) According to Theorem 12.3.4, C(PX) = C(X). And, in light of Corollary 4.2.4, C(VPx) = C(VX). Thus, Condition (d) is equivalent to Condition (a). (e) Let y represent an arbitrary vector in C(V) nC(Px). Then, y = Va for some vector a and y = Pxb for some vector b, implying (since, according to Theorem 12.3.4, Px is idempotent) that y = PxPxb = Pxy = PxVa e C(PXV). Thus, C(V)nC(Px) C C(PXV). (S.12) Now, suppose that X(X'X)~d is a solution, for every d e C(X'), to the problem of minimizing a'Va subject to X'a = d. Then, Condition (c) is satisfied (i.e., PXV = VPX), implying [since, clearly, C(PXV) c C(Px) and C(VPX) C C(V)] that C(VPX) C C(V)nC(Px) and also [in light of result (S.12)] that C(V)fiC(Px) C C(VPx). Thus, C(VPx) = C(\) n C(PX) [i.e.. Condition (e) is satisfied]. Conversely, suppose that C(VPX) = C(V)nC(Px). Then, obviously, C(VPX) C C(Px) [i.e., Condition (d) is satisfied], implying that X(X'X)~d is a solution, for every d 6 C(X'), to the problem of minimizing a'Va subject to X'a = d. (f) Since [as noted in the proof of the necessity and sufficiency of Condition (d)] C(PX) = C(X) and C(VPX) = C(VX), Condition (f) is equivalent to Condition (e). EXERCISE 12. Let V represent an n x n symmetric nonnegative definite matrix, Wan/jx n matrix, X an n x p matrix, and d a p x 1 vector. Further, let K represent any n x q matrix such that C(K) = C(\ - Px.w). Show that if rank(X'WX) = rank(X), then each of the following two conditions is necessary and sufficient for the vector WX(X'WX)~d to be a solution, for every d e C(X'), to the problem of minimizing the quadratic form a'Va (in a) subject to X'a = d: (a) V = XRIX, + (I-Px.^v')R2(I-Px.\v'), for some p x p matrix Rj and some n x n matrix R2: (b) V = XSiX, + KS2K/ for some p x p matrix Sj and some q x q matrix Si. And, show that if rank(X'WX) = rank(X) and W is nonsingular, then another necessary and sufficient condition is: (c) V = rW-1+XTiX, + KT2K' for some scalar /, some p x p matrix Tj, and some q x q matrix T2. [Hint. To establish the necessity of Condition (a), begin by observing that V = CC; for some matrix C and by expressing C as C = P\.\v'C + (1- Px.\v')C.]
19. Minimization of a Second-Degree Polynomial 219 Solution. Assume that rank(X'WX) = rank(X). Then, since X'W'X = (X'WX)', rank(X'W'X) = rank(X). Thus, applying Parts (1) and (3) of Lemma 19.5.5 (with W in place of W), we find that Px,w'X = X and Px w, = Px.w'- And« applying Part (2) of Lemma 19.5.5, we find that X'W'Px,w' = X'W' and hence that PXAV,WX = [X'W'PX.W']' = (X'W)' = WX. To establish the necesity and sufficiency of Condition (a), it suffices (in light of Theorem 19.5.1) to show that Condition (a) is equivalent to the condition that C(VWX) c C(X). If Condition (a) is satisfied, then VWX = XRjX'WX + (1- PX,W>)R2(WX ~ PX<W>WX) = XRjX'WX + (1- PX.W)R2(WX - WX) = XRjX'WX, and consequently C(VWX) C C(X). Conversely, suppose that C(VWX) C C(X). Then, VWX = XQ for some matrix Q, so that (I - Px.w')VWX = (I - Px.w)XQ = 0. (S. 13) Now, observe (in light of Corollary 14.3.8) that there exists a matrix C such that V = CC. Thus, V = [Px.w'C + (1- Px.w')C][Px.w'C + d- Px.w')C]' = px.w'CC'Px w, + Px.w'CC'(I - Px w,) +(1 - Px.w')CC'Px w, + (1- Px.w')CC'(I - Px w,). Moreover, (I ~ Px.w')vWX = OCC'WX + OCC'O +(1 - Px.w')CC'WX + (1- Px.w')CC0 = (I-Px,w0CC'WX. (S.14) Together, results (S.13) and (S.14) imply that (I-Px.w')CC'WX = 0, so that (I - Px.w)CC'Px w, = (I - Px,w)CC'WX[(X'W'X)-]'X' = 0 and Px.w'CC'(I - Px w,) = [(I - Px.w')CC'Px w,]' = 0. We conclude that V = Px.w'CC'P'xw, + (1- Px.w)CC'(I- Pxw) = XR,X' + (1- PX.W')R2(I - PX.W')',
220 19. Minimization of a Second-Degree Polynomial where Rj = (X'W'XrX'W'CC'WX[(X'W'X)-]' and R2 = CC. Thus, Condition (a) is equivalent to the condition that C(VWX) C C(X). To establish the necessity and sufficiency of Condition (b), it suffices to show that Conditions (a) and (b) are equivalent. According to Lemma 4.2.2, there exist matrices A and B such that I — Px,w' = KA and K = (I — Px,w')B. Thus, if Condition (a) is satisfied, then V = XSiX/ + KS2K/f where Si = Rj and S2 = AR2A'. Conversely, if Condition (b) is satisfied, then V = XR,X' + (1- Px.w')R2(I - Px,w0\ where Rj = Sj and R2 = BS2B'. Thus, Conditions (a) and (b) are equivalent. Assume now that W is nonsingular [and continue to assume that rank(X'WX) = rank(X)]. And [for purposes of establishing the necessity of Condition (c)] suppose that WX(X'WX)~d is a solution, for every d e C(X'). to the problem of minimizing a'Va subject to X'a = d. Then, Condition (b) is satisfied, in which case V = rW_1 + XTiX' + KT2K', where t = 0, Tj = Si, and T2 = S2. Conversely, suppose that Condition (c) is satisfied. Then, recalling that K = (I — Px.w)B for some matrix B, we find that VWX = rW_I WX + XTjX'WX + KT2K'WX = tX + X^X'WX + KTzB'd - PX$W,)WX = tX + XT,X'WX + lO^B'O = X(rI + T,X,WX). Thus, C(VWX) C C(X), and it follows from Theorem 19.5.1 that Condition (c) is sufficient (as well as necessary) for WX(X'WX)~d to be a solution, for every d e C(X'), to the problem of minimizing a'Va subject to X'a = d.
20 The Moore-Penrose Inverse EXERCISE 1. Show that, for any m x h matrix B of full column rank and for any n x p matrix C of full row rank, (BC)+ = C+B+ Solution. As a consequence of result (1.2), we have that (BC)+ = C'CCC'r^B'Br'B'. And, in light of results (2.1) and (2.2), it follows that (BC)+ = C+B+. EXERCISE 2. Show that, for any m x n matrix A, A+ = A' if and only if A'A is idempotent. Solution. Suppose that A'A is idempotent or equivalently that A'A = A'AA'A. (S.l) Then, premultiplying both sides of equality (S. 1) by A+(A+)' and postmultiplying both sides by A+, we find that A+(A+),A,AA+ = A+(A+)'A'AA'AA+ Moreover, A+(A+),A,AA+ = A+(AA+)'AA+ = A+AA+AA+ = A+AA+ = A+
222 20. The Moore-Penrose Inverse and A+(A+)'A'AA'AA+ = A+(AA+)'AA'(AA+)' = A+AA+A(AA+A)' = A+AA' = (A+A)'A' = (AA+A)' = A'. Thus,A+=A'. Conversely, suppose that A+ = A'. Then, clearly, A'A = A+A, implying (in light of Lemma 10.2.5) that A'A is idempotent. EXERCISE 3. Let T represent an m x p matrix, U an m x q matrix, V an n x p matrix, and W an n x q matrix, and define Q = W - VT~U. If C(U) C C(T) and 1Z(\) C ft(T), then /T-+T-UQ-VT- -T~UQ-\ V -Q-VT- Q- ) {*] /T U\ is a generalized inverse of the partitioned matrix I v w J, and ( Q- -Q-VT" \ V-T-UQ- T- + T-UQ-VT-; l ' i0f[u T> a generalized inverse of I _. T j. Show that if the generalized inverses T and Q" (of T and Q, respectively) are both reflexive [and if C(U) C C(T) and TZ(\) C 7£(T)], then generalized inverses (*) and (**) are also reflexive. Solution. Suppose that C(U) C C(T) and TZ(\) C TZ(T). [That these conditions are sufficient to insure that partitioned matrices (*) and (**) are generalized in- (T U\ /W V\ v w J and I f, T J, respectively, is the content of Theorem 9.6.1.] Then, TT-U = U and VT~T = V (as is evident from Lemma 9.3.5). Further, /T-+T-UQ-VT- -T-UQ-VT U\ \ -Q-VT" Q- ^V W) /T_ T_ X( -Q- _ /T~ + T-UQ-VT" -T~UQ-V _V -Q-VT- q- )\a- UQ-VT- ■vr TT- QQ)VT- QQ T-TT" + TUQ-VT-TT- I -TUQ(I-QQ-)VT- ~TUQ QQ | (S2) V-Q-VT-TT" + Q-(I - QQ-)VT~ Q-QQ_ Now, if the generalized inverses T" and Q" are both reflexive (i.e., if T"TT" = T" and Q"QQ_ = Q-), then partitioned matrix (S.2) simplifies to partitioned
20. The Moore-Penrose Inverse 223 matrix (*). We conclude that if the generalized inverses T~ and Q~ are both reflexive [and if if C(U) c C(T) and TZ(\) C ft(T)], then the generalized inverse (*) /T U\ of ( v w J [or equivalently the generalized inverse given by expression (9.6.2)] is reflexive. And, it can be shown in similar fashion that if the generalized inverses T" and Q~ are both reflexive [and if if C(U) C C(T) and TZ(\) C ft(T)], then (W V\ II T) ^0T e(luivalently the generalized inverse given by expression (9.6.3)] is reflexive. EXERCISE 4. Determine which of Penrose Conditions (1) - (4) [also known as Moore-Penrose Conditions (1)- (4)] are necessarily satisfied by a left inverse of an m x n matrix A (when a left inverse exists). Which of the Penrose conditions are necessarily satisfied by a right inverse of an m x n matrix A (when a right inverse exists)? Solution. Suppose that A has a left inverse L. Then, by definition, LA = I„. And, as previously indicated (in Section 9.2d), ALA = AI = A. Thus, L necessarily satisfies Penrose Condition (1). Further, LAL = IL = L and (LA)' = I' = I = LA, so that L also necessarily satisfies Penrose Conditions (2) and (4). However, there exist matrices that have left inverses that do not satisfy Penrose Condition (3). Suppose, for example, that A = I " J (where m > n). And, take L = (ln, K), where K is an arbitrary n x (m — n) matrix. Then, LA = I„, and AL = ( " ft I, so that L is a left inverse of A that (unless K = 0) does not satisfy Penrose Condition (3). Similarly, if A has a right inverse R, then R necessarily satisfies Penrose Conditions (1), (2), and (3). However, there exist matrices that have right inverses that do not satisfy Penrose Condition (4). EXERCISE 5. Let A represent an m x n matrix and G an n x m matrix. (a) Show that G is the Moore-Penrose inverse of A if and only if G is a minimum norm generalized inverse of A and A is a minimum norm generalized inverse of G. (b) Show that G is the Moore-Penrose inverse of A if and only if GAA' = A' andAGG' = G'. (c) Show that G is the Moore-Penrose inverse of A if and only if GA = PA> and AG = PG'. Solution, (a) By definition, G is a minimum norm generalized inverse of A if and only if AGA = A and (GA)' = GA [which are Penrose Conditions (1) and (4)], and A is a minimum norm generalized inverse of G if and only if GAG = G and (AG)' = AG [which, in the relevant context, are Penrose Conditions (2) and (3)]. Thus, G is the Moore-Penrose inverse of A if and only if G is a minimum norm
224 20. The Moore-Penrose Inverse generalized inverse of A and A is a minimum norm generalized inverse of G. (b) Part (b) follows from Part (a) upon observing (in light of Theorem 20.3.7) that G is a minimum norm generalized inverse of A if and only if GAA' = A' and that A is a minimum norm generalized inverse of G if and only if AGG' = G'. (c) Part (c) follows from Part (a) upon observing (in light of Corollary 20.3.8) that G is a minimum norm generalized inverse of A if and only if GA = PA* and that A is a minimum norm generalized inverse of G if and only if AG = PG'. EXERCISE 6. (a) Show that, for any m x n matrices A and B such that A'B = 0 andBA' = 0, (A + B)+ = A++B+ (b) Let Ai,A2,..., A& represent m x n matrices such that, for j > i = 1 k - 1, AfAy = 0 and AyAj = 0. Generalize the result of Part (a) by showing that (Ai + A2 + • • • + A*)+ = A J" + A + + • • • + A+. Solution, (a) Let X represent any n x m matrix such that (A + B)'(A + B)X = (A + B)' and Y any m x n matrix such that (A + B)(A + B)'Y = A + B. Then, since B'A = (A'B)' = 0 and AB' = (BA')' = 0, we have that (A'A + B'B)X = A' + B' and (AA' + BB')Y = A + B. Moreover, as a consequence of Corollary 12.1.2, we have that C(A) -LC(B) and [since (A')'B' = (BA')' = 0] that C(A') JLC(B'), implying (in light of Lemma 17.1.9) that C(A) fl C(B) = {0} and C(A') n C(B') = {0}. Thus, upon observing that C(A') = C(A'A), C(B') = C(B'B), C(A) = C(AA'), and C(B) = C(BB'), it follows from Theorem 18.2.7 that A'AX = A', B'BX = B', AA'Y = A, and BB'Y = B. Now, making use of Theorem 20.4.4, we find that (A + B)+ = Y'(A + B)X = Y'AX + Y'BX = A+ + B+. (b) The proof is by mathematical induction. The result of Part (b) is valid for k = 2, as is evident from Part (a). Suppose now that the result of Part (b) is valid for k = k* - 1. And, let A i A** _ j, Ak+ represent m x n matrices such that, for j > i = 1 k* — 1, AjAy = 0 and AyAj = 0. Then, observing that (Ai + • • • + A^_! )'A*. = Ai A*. + -..+ Ai*_iA*. = 0 and that AHAi + • • • + A**-i)' = A** A', + • • • + A** a;._, = 0
20. The Moore-Penrose Inverse 225 and using the result of Part (a), we find that (Ai + • • ■ + A*.-i + AH+ = KAi + ■• ■ + A*._i) + AH+ = (Ai+.-. +A*._i)+ + Aj, = A+ + ...+A+_1+A+, which establishes the validity of the result of Part (b) for k = k* and completes the induction argument. EXERCISE 7. Show that, for any m xn matrix A, (A+A)+ = A+A, and (AA+)+ = AA+. Solution. According to Corollary 20.5.2, A+A and AA+ are symmetric and idempotent. Thus, it follows from Lemma 20.2.1 that (A+A)+ = A+A and (AA+)+ = AA+. EXERCISE 8. Show that, for any n x n symmetric matrix A, AA+ = A+A. Solution. That AA+ = A+A is an immediate consequence of Part (2) of Theorem 20.5.1. Or, alternatively, this equality can be verified by making use of Part (2) of Theorem 20.5.3 (and of the very definition of the Moore-Penrose inverse). We find that AA+ = (AA+)' = (A+)'A' = (A+)'A = A+A. EXERCISE 9. Let V represent annxn symmetric nonnegative definite matrix, Xann x p matrix, and dapxl vector. Using the results of Exercises 8 and 19.11 (or otherwise), show that, for the vector X(X'X)~d to be a solution, for every d e C(X'), to the problem of minimizing the quadratic form a'Va (in a) subject to X'a = d, it is necessary and sufficient that C(V+X) C C(X). Solution. In light of the results of Exercise 19.11, it suffices to show that C(VX) C C(X) & C(V+X) C C(X). Suppose that C(VX) C C(X). Then, VX = XQ for some matrix Q. And, using the result of Exercise 8, we find that VX = VV+VX = VV+XQ = V+VXQ = V+XQ2 and hence that C(VX) C C(V+X). (S.3) Moreover, since (according to Theorem 20.5.3) V+ is symmetric and nonnegative definite, we have (in light of Lemma 14.11.2 and the result of Exercise 8) that rank(VX) > rank(V+VX) = rank(W+X) > rank(X,V+W+X) = rankCX'V+X) = rank(V+X),
226 20. The Moore-Penrose Inverse implying [since, in light of result (S.3), rank(VX) < rank(V+X)] that rank(VX) = rank(V+X). Thus, it follows from Theorem 4.4.6 that C(VX) = C(V+X). We conclude that C(V+X) C C(X). Conversely, suppose that C(V+X) C C(X). Then, V+X = XR for some matrix R. And, using the result of Exercise 8, we find that V+X = V+W+X = V+VXR = W+XR = VXR2 and hence that C(V+X) C C(VX). (S.4) Moreover, in light of Lemma 14.11.2 and the result of Exercise 8, we have that rank(V+X) > rank(W+X) = rank(V+VX) > rank(X'W+VX) = rank(X'VX) = rank(VX), implying [since, in light of result (S.4), rank(V+X) < rank(VX)] that rank(V+X) = rank(VX). Thus, it follows from Theorem 4.4.6 that C(V+X) = C(VX). We conclude that C(VX) C C(X). EXERCISE 10. Let A represent an n x n matrix. Show that if A is symmetric and positive semidefinite, then A+ is symmetric and positive semidefinite and that if A is symmetric and positive definite, then A+ is symmetric and positive definite. Do so by taking advantage of the result that if A is symmetric and nonnegative definite (and nonnull), then A+ = T+(T+)' for any matrix T of full row rank (and with n columns) such that A = T/T. Solution. Suppose that A is symmetric and nonnegative definite. Further, assume that A is nonnull — if A = 0, then A is positive semidefinite, and A+ = 0, so that A+ is also positive semidefinite (and symmetric). Then, it follows from the result cited in the exercise [which is taken from Theorem 20.4.5] that A+ = T+(T+)/ for any matrix T of full row rank (and with n columns) such that A = T/T. Thus, A+ is symmetric and (in light of Corollary 14.2.14) nonnegative definite. And, since (T+)' has n columns and since [in light of Part (1) of Theorem 20.5.1] rank (T+)' = rank T+ = rank T, it follows from Corollary 14.2.14 that A+ is positive semidefinite if rank (T) < n or equivalently if A is positive semidefinite and that A+ is positive definite if rank(T) = n or equivalently if A is positive definite. EXERCISE 11. Let C represent an m x n matrix. Show that, for any m x m idempotent matrix A, (AC)+A' = (AC)+ and that, for any n x n idempotent matrix B, B'(CB)+ = (CB)+. Solution. According to Corollary 20.5.5, (AC)+ = [(AC)'AC]+(AC)' = [(AC)'AC]+C'A\
20. The Moore-Penrose Inverse 227 and (CB)+ = (CB)'[CB(CB)']+ = B'C'[CB(CB)']+. Thus, (AC)+A' = [(AC)'AC]+C'A'A' = [(AC)'AC]+C(AA)' = [(AC)'AC]+CA' = (AC)+, and B'(CB)+ = B'B'C'[CB(CB)']+ = (BB)'C'[CB(CB)']+ = B'C'[CB(CB)']+ = (CB)+. EXERCISE 12. Let a represent annxl vector of variables, and impose on a the constraint X'a = d, where X is an n x p matrix and d a p x 1 vector such that d € C(X'). And, define /(a) = a'Va — 2b'a, where V is an n x n symmetric nonnegative definite matrix and b is an n x 1 vector such that b € C( V, X). Further, let R represent any matrix such that V = R'R, let ao represent any n x 1 vector such that X'ao = d, and take s to be any n x 1 vector such that b = Vs + Xt for some p x 1 vector t. Show that /(a) attains its minimum value (under the constraint X'a = d ) at a point a* if and only if a* = a0 + [R(I - Px)]+R(s - a0) + {I - [R(I - Px)]+R}tt - Px)w for some n x 1 vector w. Do so by, for instance, using the results of Exercise 11 in combination with the result that, for any n x k matrix Z whose columns span Af(X'), /(a) attains its minimum value (subject to the constraint X'a = d) at a point a* if and only if a* = a0 + Z(Z'VZ)-Z'(b - Va0) + Z[I - (Z'VZrZ'VZJw for some k x 1 vector w. Solution. Take Z = I - Px. Then, according to Lemma 12.5.2, C(Z) = JV(X'). And, it follows from the cited result on constrained minimization (which is taken from Section 19.6) that /(a) attains its minimum value (under the constraint X'a = d) at a point a* if and only if a* = a0 + Z(Z'VZ)+Z'(b - Va0) + [I - Z(Z,VZ)+Z,V]Zw for some n x 1 vector w. Moreover, according to Part (9) of Theorem 12.3.4, Z is symmetric and idem- potent. Thus, making use of Corollary 20.5.5 and of the results of Exercise 11, we find that Z(Z,VZ)+Z,V = Z[(RZ),RZ]+(RZ),R = Z(RZ)+R = (RZ)+R.
228 20. The Moore-Penrose Inverse And, since [in light of Part (1) of Theorem 12.3.4] Z'X = ZX = 0, Z(Z'VZ)+Z'(b - Va0) = Z(Z'VZ)+Z'(Vs + Xt - Va0) = Z(Z'VZ)+Z'V(s - ao) = (RZ)+R(s - a0). We conclude that /(a) attains its minimum value (under the constraint X'a = d) at a point a* if and only if a* = a0 + (RZ)+R(s - a0) + [I - (RZ)+R]Zw for some n x 1 vector w. EXERCISE 13. Let A represent an n x n symmetric nonnegative definite matrix, and let B represent an n x n matrix. Suppose that B — A is symmetric and non- negative definite (in which case B is symmetric and nonnegative defimte). Show that A+ — B+ is nonnegative definite if and only if rank(A) = rank(B). Do so by, for instance, using the results of Exercises 1 and 18.15, the result that W_1 — V-1 is nonnegative definite for any m x m symmetric positive defimte matrices W and V such that V — W is nonnegative defimte, and the result that the Moore-Penrose inverse H+ of a k xk symmetric nonnegative defimte matrix H equals T^T"1")', where T is any matrix of full row rank (and with k columns) such that H = T/T. Solution. Let r = rank(B). And, assume that r > 0 — if r = 0, then B = 0 and (in light of Lemma 4.2.2) A = 0, in which case A+ — B+ = 0 — 0 = 0 and rank(A) = 0 = rank(B). Then, according to Theorem 14.3.7, there exists anrxn matrix P such that B = I^P. Similarly, according to Corollary 14.3.8, there exists a matrix Q such that A = Q'Q. And, according to the result of Exercise 18.15, 11(A) C 11(B), (S.5) implying [since 1Z(Q) = 11(A) and 1Z(P) = 11(B)] that 1l(Q) C ft(P) and hence that there exists a matrix K (having r columns) such that Q = KP. Thus, B-A = P'P-Q'Q = P'(I-K/K)P. Moreover, according to Lemma 8.1.1, P has a right inverse R, so that I - K'K = (PR)'(I - K'K)PR = R'(B - A)R. And, as a consequence, I — K'K is nonnegative definite. Now, suppose that rank(A) = rank(B) (= r). Then, /■ = rank(Q) = rank(KP) < rank(K),
20. The Moore-Penrose Inverse 229 implying (since clearly rank K < r) that rank(K) = r and hence (in light of Corollary 14.2.14) that K'K is positive definite. Thus, it follows from one of the cited results (a result encompassed in Theorem 18.3.4) that (K'K)_1 - I is nonnegative definite. Moreover, upon observing that A = P'CK'KJP, it follows from the result of Exercise 1 that A+ = P+ (K'K)-1 {V)+ = P+(K/K)~1(P+)' and from another of the cited results (a result covered by Theorem 20.4.5) that B+ = P+(P+)', so that A+ - B+ = P+KK'K)-1 - I](P+)'. And, in light of Theorem 14.2.9, we conclude that A+ —B+ is nonnegative definite. Conversely, suppose that A+—B+ is nonnegative definite. Then, it follows from the result of Exercise 18.15 that ft(B+) C ft(A+), implying that rank(B+) < rank(A+) and hence [in light of Part (1) of Theorem 20.5.1] that rank(B) < rank(A). Since [in light of result (S.5)] rank(A) < rank(B), we conclude that rank(A) = rank(B).
21 Eigenvalues and Eigenvectors EXERCISE 1. Show that an /2 x n skew-symmetric matrix A has no nonzero eigenvalues. Solution. Let X represent any eigenvalue of A and let x represent an eigenvector that corresponds to X. Then, —A'x = Ax = Ax, implying that -A'Ax = -A'(Xx) = A(-A'x) = X(Xx) = X2x and hence that —x'A'Ax = X2x/x. Thus, observing that x^O and that A'A is nonnegative definite, we find that 0 < X2 = -x/A,Ax/x,x < 0, leading to the conclusion that X2 = 0 or equivalently that X = 0. EXERCISE 2. Let A represent annxn matrix, B a k x k matrix, and X an n x k matrix such that AX = XB. (a) Show that C(X) is an invariant subspace (of TZ"xl) relative to A. (b) Show that if X is of full column rank, then every eigenvalue of B is an eigenvalue of A. Solution, (a) Corresponding to any (n x 1) vector u in C(X), there exists a k x 1 vector r such that u = Xr, so that Au = AXr = XBr € C(X). Thus, C(X) is an invariant subspace relative to A.
232 21. Eigenvalues and Eigenvectors (b) Let X represent an eigenvalue of B, and let y represent an eigenvector of B corresponding to X. By definition. By = Xy, so that A(Xy) = XBy = X(Xy) = X(Xy). Now, suppose that X is of full column rank. Then (since y ^ 0) Xy £ 0, leading us to conclude that X is an eigenvalue of A (and that Xy is an eigenvector of A corresponding to X). EXERCISE 3. Let p(X) represent the characteristic polynomial of an n x n matrix A, and let <?o, c\, c2 c„ represent the respective coefficients of the characteristic polynomial, so that ii p(X) = c0X° + ciX + c2X2 + • • • + cn\" = J2 cs*S s=0 (for X e 11). Further, let P represent the n x n matrix obtained from p(X) by formally replacing the scalar X with the n x n matrix A (and by setting A0 = I,,). That is, let n P = c0I + ciA + c2A2 + ... + c,JA" = £]c,A*. 5=0 Show that P = 0 (a result that is known as the Cayley-Hamilton theorem) by carrying out the following four steps. (a) Letting B(X) = A -XI and letting H(X) represent the adjoint matrix of B(X), show that (for X e 1Z) H(X) = Ko + XK, + X2K2 + • • • + X^X-i , where Ko, Kj, K2 K„_i are n x n matrices (that do not vary with X). (b) Letting T0 = AKo, T„ = -K„_i, and (for s = 1 n - 1) Ts = AK* - Ks-1, show that (for X e 11) To + XTj +X2T2 + ■■ • + X"T„ = p(X)I„ . [Hint. It follows from a fundamental result on adjoint matrices that (for X e 1Z) B(X)H(X) = |B(X)H, = p(X)I„.] (c) Show that, for s = 0, 1 n, T.v = c5I. (d) Show that P = To + ATi + A2T2 + • - • + A" Tw = 0. Solution, (a) Let /?;y(X) represent the //th element of H(X). Then, /i;y(X) is the cofactor of the y/th element of B(X). And, it is apparent from the definition of a
21. Eigenvalues and Eigenvectors 233 cofactor and from the definition of a determinant [given by formula (13.1.2)] that hjj{X) is a polynomial (in X) of degree n - 1 or n - 2. Thus, huM = C + $h + k$x2 + • • • + fc'rV-1 for some scalars *{?}, fcjj )%k\j) kj"~l ] (that do not vary with X). And, it follows that H(X) = K0 + XKi + X2K2 + • • • + X""1^-!, where (for s = 0,1,2,..., n — 1) K5 is the n x /i matrix whose ijth element is kij ' (b) In light of Part (a), we have that (for X e 1Z) B(X)H(X) = (A - XI)(Ko + XK, + X2K2 + • • • + X""1^-,) = T0 + XT1+X2T2 + --- + X"Tn. And, making use of the hint, we find that (for X e 1Z) To + XTi + X2T2 + • • • + X"T„ = p(X)I„ . (c) For s = 0, 1 /i, let tjp represent the ijth element of Ts. Then, it follows from Part (b) that (for X e 1Z) ,(0),,,(1),,^(2), _i_ynt(n)_ ( P(X), ifj = L hi +xtij +Ar»7 +••• + * ty -| o, if7#/. Consequently, As) _ f cs, if j = /, 'v ~\ 0, if j #/, and hence Ts = csI. (d) Making use of Part (c), we find that T0 + AT!+A2T2 + . ■•H-A'X = c0I + A(c,I) + A2(c2I) + • • • + A"(cD = P. Moreover, T0 + ATi+A2T2 + --. + A"T„ = (A-A)Ko + (A-A)AK, + (A-A)A2K2 + ... +(A-A)A"-|Kll-i = 0. EXERCISE 4. Let co, ci cn-\* c„ represent the respective coefficients of the characteristic polynomial p(X) of an n x n matrix A [so that p(X) = Co +
234 21. Eigenvalues and Eigenvectors ci A + • • • + c„_iX"-1 + cnXn (for X € ft)]. Using the result of Exercise 3 (the Cayley-Hamilton theorem), show that if A is nonsingular, then cq # 0, and A"1 = -(l/co)(ciI + c2A +... + C.A"-1). Solution. According to result (1.8), cq = |A|, and, according to the result of Exercise 3, c0I + ci A + c2A2 + • • • + c„A" = 0. (S.l) Now, suppose that A is nonsingular. Then, it follows from Theorem 13.3.7 that co ^ 0. Moreover, upon premultiplying both sides of equality (S.l) by A~l, we find that c0A_1 + cil + c2A + c3A2 + • • • + ChA"-1 = 0 and hence that A"1 = Hl/c0)(ciI + c2A + c3A2+ --- + ^-1). EXERCISE 5. Show that if an n x n matrix B is similar to an n x n matrix A, then (1) B* is similar to A* (k = 2,3,...) and (2) B' is similar to A'. Solution. Suppose that B is similar to A. Then, there exists an n x n nonsingular matrix C such that B = C"1 AC. (1) Clearly, it suffices to show that (for k = 1,2,3,...) B*r = C-1A*C. Let us proceed by mathematical induction. Obviously, B1 =ClA.lC. Now, suppose that B*_I = CT1 A*_IC (where k > 2). Then, B* = BBA_1 = C"1 ACC"1 A*_IC = C_IA*C. (2) We find that b' = (C-^c/ = CA'tc-1/ = [(Cr'r'A'tc'r1. Thus, B' is similar to A;. EXERCISE 6. Show that if an n x n matrix B is similar to an (/? x n) idempotent matrix, then B is idempotent. Solution. Let A represent an n x n idempotent matrix, and suppose that B is similar to A. Then, there exists an n x n nonsingular matrix C such that B = C"1 AC. And, it follows that B2 = C1 ACC-'AC = C1 A2C = C1 AC = B. EXERCISE 7. Let A = f Q " J and B = ll J J. Show that B has the same rank, determinant, trace, and characteristic polynomial as A, but that, nevertheless, B is not similar to A.
21. Eigenvalues and Eigenvectors 235 Solution. Clearly, |B| = 1 = |A|, rank(B) = 2 = rank(A), tr(B) = 2 = tr(A), and the characteristic polynomial of both B and A is p(k) = (X - I )2. Now, suppose that CB = AC for some 2x2 matrix C = {<?,;}. Then, since CB=/c„ cn+cn\ and AC = c=(cn cn\ \C2l C21+C22/ \C2J C22/ c\ 1 + C12 = C12 and C21 + C22 = C22. implying that c\ \ = 0 and C21 = 0 and hence that C is singular. Thus, there exists no 2 x 2 nonsingular matrix C such that CB = AC. And, we conclude that B is not similar to A. EXERCISE 8. Expand on the result of Exercise 7 by showing (for an arbitrary positive integer n) that for an n x n matrix B to be similar to an n x n matrix A it is not sufficient for B to have the same rank, determinant, trace, and characteristic polynomial as A. Solution. Suppose that A = In, and suppose that B is a triangular matrix, all of whose diagonal elements equal 1. Then, in light of Lemma 13.1.1 and Corollary 8.5.6, B has the same determinant, rank, trace, and characteristic polynomial as A. However, for B to be similar to A, it is necessary (and sufficient) that there exist an n x n nonsingular matrix C such that B = C_1 AC or equivalently (since A = In) that B = I„. Thus, it is only in the special case where B = I„ that B is similar to A. EXERCISE 9. Let A represent an n x n matrix, B a k x k matrix, and X an n x k matrix such that AX = XB. Show that if X is of full column rank, then there exists an orthogonal matrix Q such that Q'AQ = (n H T12), where Tn is a k x k matrix that is similar to B. Solution. Suppose that rank(X) = k. Then, according to Theorem 6.4.3, there exists an n x k matrix U whose columns are orthonormal (with respect to the usual inner product) vectors that form a basis for C(X). And, X = UC for some kxk matrix C. Moreover, C is nonsingular [since rank(X) = k]. Now, observe that AUC = AX = XB = UCB and hence that AU = AUCC"1 = UCBC-1. Then, in light of Theorem 21.3.2, there exists an orthogonal matrix Q such that Q'AQ = (Q n J}2), where Tn = CBC"1. Moreover, Tn is similar toB. EXERCISE 10. Show that if 0 is an eigenvalue of an n x n matrix A, then its algebraic multiplicity is greater than or equal to n — rank(A). Solution. Suppose that 0 is an eigenvalue of A. Then, according to Theorem 21.3.4, its algebraic multiplicity is greater than or equal to its geometric multiplicity, and, according to Lemma 11.3.1, its geometric multiplicity equals w — rank(A). Thus,
236 21. Eigenvalues and Eigenvectors the algebraic multiplicity of the eigenvalue 0 is greater than or equal to n—rank(A). EXERCISE 11. Let A represent annxn matrix. Show that if a scalar X is an eigenvalue of A of algebraic multiplicity y, then rank(A — XI) > n — y. Solution. Suppose that X is an eigenvalue of A of algebraic multiplicity y. Then, making use of Theorem 21.3.4 and result (1.1) {and recalling that, by definition, the geometric multiplicity of X equals dim[jV(A — XI)]}, we find that y > dimt^A - XI)] = n - rank(A - XT) and hence that rank(A — XI) > n — y. EXERCISE 12. Let y\ represent the algebraic multiplicity and v\ the geometric multiplicity of 0 when 0 is regarded as an eigenvalue of an n x n (singular) matrix A. And let yi represent the algebraic multiplicity and 1¾ the geometric multiplicity of 0 when 0 is regarded as an eigenvalue of A2. Show that if vi = yi, then V2 = n = vi. Solution. For any n x 1 vector x such that Ax = 0, we find that A2x = AAx = A0 = 0. Thus, V2 = dim[MA2)] > dimt^A)] = vi. (S.2) There exists an n x v\ matrix U whose columns form an orthonormal (with respect to the usual inner product) basis for jV(A). Then, AU = 0 = U0. And, it follows from Theorem 21.3.2 that there exists an n x (n — vi) matrix V such that the nxn matrix (U, V) is orthogonal and, taking V to be any such matrix, that (U,V)'A(U,V)=(J ^) n V A V) J* M°reover' it follows from Theorem 21.3.1 that y\ equals the algebraic multiplicity of 0 when 0 is regarded as an eigenvalue ft V'AV) and hence (in light of Lemma 21.2.1) that y\ equals v\ plus the algebraic multiplicity of 0 when 0 is regarded as an eigenvalue of V'AV. Now, suppose that v\ = y\. Then, the algebraic multiplicity of 0 when 0 is regarded as an eigenvalue of V'AV equals 0; that is, 0 is not an eigenvalue of V'AV. Thus, it follows from Lemma 11.3.1 that V'AV is nonsingular. Further, (U, V)'A2(U, V) = (U, V)'A(U, V)(U, V)'A(U, V) _/0 U'AV\2_/0 U'AVVAVN " \0 YA\) ~ \p (VAV)2 y 4
21. Eigenvalues and Eigenvectors 237 o /0 U'AW'AV\ so that Az is similar to I , 2 J • And» since (V'AV)2 is nonsingular, it follows from Lemma 21.2.1 that yi = v\. Recalling inequality (S.2), we conclude, on the basis of Theorem 21.3.4, that v\ = yi > V2 > v\ and hence that 1¾ = Yi = vi. EXERCISE 13. Let Xi and X2 represent eigenvectors of an n x n matrix A, and let C\ and C2 represent nonzero scalars. Under what circumstances is the vector x = cjXi + C2X2 an eigenvector of A? Solution. Let k\ and X2 represent the two (not-necessarily-distinct) eigenvalues to which xi and X2 correspond. Then, by definition, Axi = X1X1 and AX2 = A2X2, and consequently AX = C!AXi + C2AX2 = C1X1X1 + C2*2*2 = *]X + (*2 - Al)C2*2 ■ Thus, if A2 = Xi. then x is an eigenvector of A [unless X2 = -(ci/C2>xi, in which case x = 0]. Alternatively, if X2 # Ai, then, according to Theorem 21.4.1, xi and X2 are linearly independent, implying that x and X2 are linearly independent (as can be easily verified) and hence that there does not exist any scalar c such that Xix + (A2 - *i)QX2 = cx- We conclude that if X2 # Ai then x is not an eigenvector of A. EXERCISE 14. Let A represent annxn matrix, and suppose that there exists an n x n nonsingular matrix Q such that Q_1 AQ = D for some diagonal matrix D = [di). Further, for 1 = 1 «, let rj represent the /th row of Q_1. Show (a) that A' is diagonalized by (Q-1)', (b) that the diagonal elements of D are the (not necessarily distinct) eigenvalues of A', and (c) that ri 17, are eigenvectors of A' (with r/ corresponding to the eigenvalue di). Solution. The validity of Part (a) is evident upon observing that D = Dx = (CT'AQ)' = Q'A^Q"1)' = [(Q'r'r^CT1)' = i(Q-l)TlA!(Q-ly. And, observing also that 17 is the /th column of (Q-1)\ the validity of Parts (b) and (c) follows from Parts (7) and (8) of Theorem 21.5.1. EXERCISE 15. Show that if an n x n nonsingular matrix A is diagonalized by annxn nonsingular matrix Q, then A-1 is also diagonalized by Q. Solution. Suppose that A is diagonalized by Q. Then, it follows from Theorem 21.5.1 that the columns of Q are eigenvectors of A and hence (in light of Lemma 21.1.3) of A-1, leading us to conclude (on the basis of Theorem 21.5.2) that Q diagonalizes A-1 as well as A. [Another way to see that Q diagonalizes A-1 is to observe that Q-1A_1Q = (Q-1AQ)_1 and that (since Q_1AQ is a diagonal matrix) (Q_1AQ)_1 is a diagonal matrix.]
238 21. Eigenvalues and Eigenvectors EXERCISE 16. Let A represent an n x n matrix whose spectrum comprises k eigenvalues X\ A* with algebraic multiplicities y\ yjs respectively, that sum to n. Show that A is diagonalizable if and only if, for i = 1,..., k, rank(A — A/I) = n — y/. Solution. Let vi v* represent the geometric multiplicities of Ai A*, respectively. Then, according to result (1.1), v,- = n — rank(A — A/1) (/ = 1 k). Thus, rank(A — A,-1) = n — yi -& Yi=n — rank(A — A/I) ^ v,- = yi. Moreover, it follows from Corollary 21.3.7 that v,- = v/ for i = 1,..., k if and only if £j_, v,- = £}=1 yj or equivalently (since £j=1 yj = n) if and only if Y$-i Vi = '*• We conclude, on the basis of Corollary 21.5.4, that A is diagonalizable if and only if, for i = 1 k% rank(A — A,-1) = n — y,-. EXERCISE 17. Let A represent an n x n symmetric matrix with not-necessarily- distinct eigenvalues d\ d„ that have been ordered so that d\ < <h < • • • < d„. And, let Q represent an n x n orthogonal matrix such that Q'AQ = diag(di dn) — the existence of which is guaranteed. Further, for m = 2 n — 1, define Sm = [xeKnxl : x #0, Q;„x = 0} and rm = {x e 1Znxl : x#0, Pjttx = 0}, where Qm = (qL q,,,^) and P,„ = (q,,l+1 q„). Show that, for m = 2,...,/2-1, . x'Ax x'Ax dm = mm —— = max —— . xeSm x'x xeTm x'x Solution. Let x represent an arbitrary n x 1 vector, and let y = Q'x. Partition Q and y as Q = (Q,„, R,«) and y = I J1 I (where y, has m — 1 elements). Then, yi=QmX.V2 = R>>and x = Qy = Qwy,+Rffly2. Moreover, since the columns of Rm are linearly independent and the columns of Q,„ are linearly independent, y2 = 0 ^ Rmy2 = 0 (or equivalently y2 # 0 «£► R»y2 * 0) and y, = 0 & Qmy, = 0. Thus, " Q,'„x = 0 ^ y,=0 ^ x = Rmy2. It follows that x e Snt if and only if x = Rmy2 for some (n - m + 1) x 1 nonnull vector y2. It is now clear (since R,'nR,„ = I) that . x'Ax . (Rmy2)'ARwy2 . y;(R,'„AR,„)y.> mm—- = mm— — = mm-——, -. *esm x'x y254o (Rwy2)'RIMy2 y254o y;y2 Moreover, R^AR,,, = diag(dIM, </,„+I d„). so that R,'MARIM is a symmetric matrix whose smallest eigenvalue is dm . Thus, as a consequence of Theorem
21. Eigenvalues and Eigenvectors 239 21.5.6, we have that _. _y'2(R;HARm)y2 mm ; = d,„ . V2^ y2y2 And, we conclude that . x'Ax mm —— = dm . xeS,„ x'X That maxxerwl x'Ax/x'x = dm follows from a similar argument. EXERCISE 18. Let A represent an n x n symmetric matrix, and adopt the following notation: d\ d„ are the not-necessarily-distinct eigenvalues of A, qi q„ are orthonormal eigenvectors corresponding to d\ d„t respectively, Q = (qj q,,), [X\ Xk) is the spectrum of A; and, for j = 1 k, Sj = [i : d{ = Xjl Ej = "£ieSj q,q}t and Qy = (q,-, q,-,,.), where i\ iv. denote the elements of Sj. (a) Show that the matrices Ei Ek, which appear in the spectral decomposition A = X)y=i \/Ey « have the following properties: (1) Ei+--. + E* = I; (2) Ei Ek are nonnull, symmetric, and idempotent; (3) for/#; = 1,...,A:, EtEj=0; and (4) rank(Ei) + • • • + rank(E*) = n. (b) Take Fj,..., Fr to be n x n nonnull idempotent matrices such that Fi + h Fr = I. And, suppose that, for some distinct scalars tj t>, A = TlFi-|-----|-TrFr. Show that r = k and that there exists a permutation t\,..., tr of the first r positive integers such that (for j = \ r)Tj =Xtj and Fy = E^. Solution, (a)(1) k n y=i teSj /=1 (2) Observe that (for ; = 1,..., k) Ej = QyQ} and Q^Qy = I. Then, clearly, Ey is symmetric. And, Qy is nonnull, implying (in light of Corollary 5.3.2) that Ey is nonnull. Further Ej = QyQyQyQy = QylQy = QyQy = Ey , so that Ey is idempotent. (3) and (4) In light of Theorem 18.4.1, Parts (3) and (4) follow from Parts (1) and (2).
240 21. Eigenvalues and Eigenvectors (b) Observe (in light of Theorem 18.4.1) that, for t ^ j = 1 r, F,Fy = 0. Then, for j = 1,..., r, we find that AFy = nFiF; + • • • + rrFrFy = xjtfj = ryFy , implying that Ty is an eigenvalue of A and that any nonnull column of Fy is an eigenvector of A corresponding to Ty . Consequently, there exists some subset T = {t\ tr) of the first k positive integers such that (for j = 1 r) Tj = XtJ (so that r < k). Further, C(EtJ) = C(Qr.Q;.) = C(Qtj) = MA - A,,D = JV(A - Tyl), so that Fy = E,.Ly for some n x n matrix Ly . We have that A = ^,^,^ + --- + ^,^^, implying [in light of Part (a) and equality (5.5)] that (for ; = 1 r) ^-tj^tj = ^ijEr = E/;A = XfjEfjhj = XtJEtjhj = X/yFy , (S.3) so that Xtj = 0 or F; = Etj. Moreover, for j ¢. 7\ we find that XjEj = XjE) = EyA = 0, implying (since Ey ^ 0) that Xy = 0 (and hence that k < r + 1). To complete the proof, it suffices to show that r = k and that (for j = 1,..., r) Fy = Etj. Let us consider separately the following two cases: (1) Xtj ¥" 0 for y = l r; and (2) Xts = 0 for some integer s (1 < s < r). In Case (1), it follows from result (S.3) that (for j = 1 r) Fy = E,r Moreover, r = k, since otherwise there would exist a positive integer ^¢^ such that Xs = 0, and we would have [in light of Part (a)] that E, = I-^E0=I-^Fy =1-1 = 0, which [since, according to Part (a), Es ^ 0] would lead to a contradiction. In Case (2), it is clear that r = k and also [in light of result (S.3) and Part (a)] that Fy = Etj for j ^ s and EXERCISE 19. Let A represent anwxn symmetric matrix, and let d\ d„ represent the (not-necessarily-distinct) eigenvalues of A. Show that Hitu—oo A* = 0 if and only if, for / = 1 «, \d{\ < 1.
21. Eigenvalues and Eigenvectors 241 Solution. Let D = diag(d\ dn). Then, according to Corollary 21.5.9, there exists ann x n orthogonal matrix Q such that Q'AQ = D. And, it follows from result (5.6) that A* = QD*Q'. Moreover, since D* = diag(df d,f), 141 < 1 for / = 1 n <& lim d:k =0 for / = 1 n *—CO <& lim D* = 0. Now, suppose that, for i = 1 n, \di\ < 1. Then, lim A* = Q( lim D*)Q' = QOQ7 = 0. A-*oo £-»-oo Conversely, suppose that lim/^oo A* = 0. Then, observing that D* = Q'A*Q, we find that lim D* = lim Q'A*Q = Q'( lim A*)Q = tfOQ = 0. k-KX, k—KX> k-►oo And, it follows that, for i = 1,...,w, \dj\ < 1. EXERCISE 20. (a) Show that if 0 is an eigenvalue of an n x n not-necessarily- symmetric matrix A, then it is also an eigenvalue of A+ and that the geometric multiplicity of 0 is the same when it is regarded as an eigenvalue of A+ as when it is regarded as an eigenvalue of A. (b) Show (via an example) that the reciprocals of the nonzero eigenvalues of a square nonsymmetric matrix A are not necessarily eigenvalues of A+. Solution, (a) Suppose that 0 is an eigenvalue of A. Then, according to Lemma 21.1.1, rank(A) < n. And, since (according to Theorem 20.5.1) rank(A+) = rank(A), it follows that rank(A+) < n, implying (in light of Lemma 21.1.1) that 0 is also an eigenvalue of A+. Further, making use of Lemma 11.3.1, we find that dimfMA"1")] =/2- rank(A+) = n - rank(A) = dimfMA)], so that the geometric multiplicity of 0 is the same when it is regarded as an eigenvalue of A+ as when it is regarded as an eigenvalue of A. (b) Consider the n x n matrix A = (1,,, 0) (where n > 2). We find that A'A = diag(«, 0,0 0) and hence that (A'A)+ = diag(l//i, 0,0 0), implying (in light of Corollary 20.5.5) that A+ = (^)^=(0^). Since A and A+ are triangular, it is easy to see that the distinct eigenvalues of A are 0 and 1, while those of A+ are 0 and l/n. Obviously, 1//2 is not (for n > 2) the reciprocal of 1.
242 21. Eigenvalues and Eigenvectors EXERCISE 21. Show that, for any positive integer n that is divisible by 2, there exists an n x n orthogonal matrix that has no eigenvalues. [Hint. Find a 2 x 2 orthogonal matrix Q that has no eigenvalues, and then consider the block-diagonal matrix diag(Q, Q Q).] Solution. Let Q = ( . ft J. Clearly, Q is orthogonal, and (as shown in Section 21.1) it has no eigenvalues. Now, consider the n x n block-diagonal matrix diag(Q, Q Q) (having n/2 diagonal blocks). This matrix is orthogonal (as is easily verified), and, as a consequence of Part (2) of Lemma 21.2.1, it has no eigenvalues. EXERCISE 22. Let Q represent an n x n orthogonal matrix, and let p(X) represent the characteristic polynomial of Q. Show that (for X ^ 0) p(X) = ±X>(1A). Solution. Let X represent an arbitrary nonzero scalar. Then, Q - XI = Q - XQQ' = -XQ[Q' - (1/X)I] = -XQ[Q - (1/X)I]'. Thus, making use of Theorem 13.3.4, Lemma 13.2.1, and Corollaries 13.2.4 and 13.3.6, we find that Pfr) = IQ ~ AI| = | - XQ||[Q - (1/X)in = (-A)"|QIIQ-(lA)I| = (-i)nx"(±i)/7(i/X) = ±x"pdA). EXERCISE 23. Let A represent an n x n matrix, and suppose that the scalar 1 is an eigenvalue of A of geometric multiplicity v. Show that v < rank(A) and that if v = rank(A), then A is idempotent. Solution. That v < rank(A) is an immediate consequence of Corollary 21.3.8. Now, suppose that v = rank(A). Then, it follows from Corollary 21.3.8 that A has no nonzero eigenvalues other than 1, and it follows from Corollary 21.5.4 that A is diagonalizable. Thus, as a consequence of Theorem 21.8.3, we have that A is idempotent. EXERCISE24. Let A represent annxn nonsingular matrix. And, let X represent an eigenvalue of A, and x represent an eigenvector of A corresponding to X. Show that |A|/X is an eigenvalue of adj(A) and that x is an eigenvector of adj(A) corresponding to |A|/X. Solution. Note (in light of Lemma 21.1.1 and Theorem 13.3.7) that X ^ 0 and |A|^0. According to Lemma 21.1.3, 1 /X is an eigenvalue of A \ and x is an eigenvector of A-1 corresponding to 1/X. And, since (according to Corollary 13.5.4) adj(A) =
21. Eigenvalues and Eigenvectors 243 |A|A \ it follows from the results of Section 21.10 that | A|/X is an eigenvalue of adj(A) and that x is an eigenvector of adj(A) corresponding to |A|/X. EXERCISE 25. Let A represent an n x n matrix, and let /?(X) represent the characteristic polynomial of A. And, let Xi X* represent the distinct eigenvalues of A, and y\ y* represent their respective algebraic multiplicities, so that (for allX) k p(X) = (-l)"^(X)f](X-Xy)^ /=1 for some polynomial q (X) (of degree n — J^j-. i Yj) that has no real roots. Further, define B = A — XjUV, where U = (uj uVl)isann x y\ matrix whose columns uj u,,, are (not necessarily linearly independent) eigenvectors of A corresponding to Xi and where V = (vj \Yx) is an n x y\ matrix such that V'U is diagonal. (a) Show that the characteristic polynomial, say r(X), of B is such that (for all n k r(X) = (-l),,^(X)f][X-(l-v;.u/)X,]f](X-Xy)^ . (E.1) ,=1 y=2 [Hint. Since the left and right sides of equality (E.1) are polynomials (in X), it suffices to show that they are equal for all X other than Xi X*.] (b) Show that in the special case where U'V = cl for some nonzero scalar c, the distinct eigenvalues of B are either X2 X5_i, X5, ks+\ X* with algebraic multiplicities yi Ys-i* Ys + Yu Ys+i»• • • * W. respectively, or (1 — c)Xi, X2, ..., Xjt with algebraic multiplicities y\, y> H-, respectively [depending on whether or not (1 — c)Xi = X* for some s (2 < s < k)]. (c) Show that, in the special case where y\ = 1, (1) uj is an eigenvector of B corresponding to the eigenvalue (1 — VjUj )Xi and (2) for any eigenvector x of B corresponding to an eigenvalue X [other than (1 — v'jUOXi], the vector x-Xiai-XrVjXjui is an eigenvector of A corresponding to X. Solution, (a) Let X represent any scalar other than Xj X*. And, observe that (A - XI)U = AU - XU = (Xj - X)U, so that (A - XI)_1U = -(X - Xi)_1U. Then, making use of Corollary 18.1.2, we find that r(X) = IA-XI-XjUV'I = IA - XI| |I - Xi (A - XD-^UV'I
244 21. Eigenvalues and Eigenvectors = p(k)\ln+Xl(k-Xl)-lV\'\ = p(x)\iYl+xl(x-xlrl\'v\ = pwnn+xi(x-xir,vs«,] 1=1 = P(k)(x - Xi)-y* Y[ (* - *i + *nfa> i=i n k = (-i)^a) f] ^ - n - vju,)^,] Y\a - kSyi. ,=i y=2 (b) Part (b) follows from Part (a) upon observing that, in this special case, I~I [X - (1 - vJu^A,] = [X - (1 - c)A,F». i=i (c) (1) In this special case, Buj =Auj -Xjuiv^uj =XjUi -Xi^uOuj =(1 -vJuOXiUi. (2) Suppose that y\ = 1, and let d = X\(X\ - A)_1vix. Since (by definition) Bx = Ax, we have that (A - Aujv'jjx = (A - Aiuiv^x = Ax and hence that A[x- (v;,x)ui] =Ax. Thus, A(x -du\) = A[x - (vix)m + (vjX)uj - du\] = Ax - (d - vJx)AiUi. Moreover, X\d - Xd = (Ai - X)d = Aiv^x, implying that {d - VX\)X\ = Xd. And, it follows that A(x - dui) = A(x - dui). Since \ — du\ ^ 0 (as is evident upon observing that x and uj are eigenvectors of B that correspond to different eigenvalues and hence, in light of Theorem 21.4.1, that x and uj are linearly independent), we conclude that x — du j is an eigenvector of A corresponding to A. EXERCISE 26. Let A represent an m x n matrix of rank r. And, take P to be any m x m orthogonal matrix and Dj to be any /• x r nonsingular diagonal matrix such that
21. Eigenvalues and Eigenvectors 245 Further, partition P as P = (Pi,P2>, where Pj has r columns, and let Q = (Qi. Q2)' where Qj = A'PjDj"1 and where Q2 is any n x (n — r) matrix such that QiQ2 = 0. Show that '«"-(? !)• Solution. Upon applying Theorem 21.12.1 with A' in place of A (and with n and m in place of m and iu respectively) and writing P = (Pj, P2) for Q = (Qi, Q2) and Q = (Q{, Q2) for P = (Pi, P2), we find that Q'A »-(?:) Thus, F-AQ-WTAV-^ *)'-(* J). Or, alternatively, the equality P'AQ = ( J ft I can be established via an argument paralleling the proof of Theorem 21.12.1. EXERCISE 27. Let A represent anmxn matrix. And, let P represent an m x m orthogonal matrix, Q an n x n orthogonal matrix, and Dj = [si] an r x r diagonal matrix with (strictly) positive diagonal elements such that P'AQ = ( J ft); denote by pj pr the first through rth columns of P and by qt q,. the first through rth columns of Q; and define ol\ ajt to be the distinct values represented among s\ sr. Then, the singular value decomposition of A is A = £*=i «yUy , where (for j = 1,..., k) Vj = *£ieLj p,q; with Lj = {i : Si =<xj}. Show that the matrices Ui Ujt, which appear in the singular value decomposition, are such that UyU^U/ = U/ (for j = 1 k) and UJU/ = 0 and V,\fj = 0 (for t # ; = 1 k). Solution. By definition, , , f 1, fori; = / PuP/=Quq, = J a forw ^,-. Thus, uyu; =(IZ p^)' IZ P'^ = IZ IZ QuPuPi-q; = 5] q,q!. veLj ieLj veLj ieLj ieLj so that veLj ieLj veLj ieLj ieLj
246 21. Eigenvalues and Eigenvectors And, for / # j, vel# ieLj veL, ieLj and u^ = £ m1<E m^ = E E Mtorf =°- i»€Lf /gL; vel, /gjL/ EXERCISE 28. Let A represent an m x n matrix. And, as in Exercise 27, take P to be an m x m orthogonal matrix, Q an n x /i orthogonal matrix, and Di an r x r nonsingular diagonal matrix such that P'AQ = I a ! a I • Further, partition P and Q as P = (Pj, P2) and Q = (Qj, Q2), where each of the matrices Pj and Q, has /• columns. Show that C(A) = C(Pj) and that J\f(\) = C(Q2). Solution. Clearly, A = p(?' !)*-**<«■ implying that C(A) c C(P\). Moreover, as a consequence of result (12.12), C(Pj) C C(A). Thus, C(A) = C(Pj). We have that ® &)-<*-«■-& U- so thatQ',Q2=0. Thus, AQ2 = PiD,Q',Q2=0. Moreover, making use of result (12.9), we find that rank(Q2) = n — r = n — rank(A). And, we conclude, on the basis of Lemma 11.4.1, that .AAA) = CCQt). EXERCISE 29. Let Aj A* represent n x n not-necessarily-symmetric matrices, each of which is diagonalizable. Show that if Aj Ajt commute in pairs, then Aj Ajt are simultaneously diagonalizable. Solution. The proof is a modified version of the mathematical induction argument employed in Section 13 (in showing that symmetric matrices that commute in pairs are simultaneously diagonalizable). That one diagonalizable matrix Aj is "simultaneously" diagonalizable is obvious. Suppose now that any k - I diagonalizable matrices (of the same order) that commute in pairs can be simultaneously diagonalized (where k > 2). And,
21. Eigenvalues and Eigenvectors 247 let Ai A* represent k diagonalizable matrices of arbitrary order n that commute in pairs. Then, to complete the induction argument, it suffices to show that Ai A* can be simultaneously diagonalized. Let Xi Xr represent the distinct eigenvalues of Ajt, and let vi,..., vr represent their respective geometric multiplicities. Take Qy to be an n x Vj matrix whose columns are linearly independent eigenvectors of A* corresponding to the eigenvalue Ay or, equivalently, whose columns form a basis for the vy-dimensional linear space Af{An — Ay I) (./ = 1 r). And, define Q = (Qj Qr). Since A* is diagonalizable, it follows from Corollary 21.5.4 that J2y=1 Vj = n, so that (in light of Theorem 21.4.2) Q has n linearly independent columns and hence is nonsingular. Further, Q-%.Q = diag(XiIV| ArIlv) (S.4) (as is evident from Theorems 21.5.2 and 21.5.1). As in the case where Ai A* are symmetric, there exists a vy x vy matrix A,Qy = QyB/y (S.5) (/ = 1 k—U j = 1 /), and we find that (for i = 1 k- 1) Q"1 A/Q = diag(B/i Bir). (S.6) Now, partition Q_l as \Tr) where (for j = 1 r) Ty has vy rows. And, note that r,Q„, = |J;- for «1 = j = 1, for/77 ^ j = 1, Then, using result (S.5), we find that, for j = 1 r, B/y = IlvB,y = TjQjBij = TyA/Qy (i = 1 it- 1) and that BsjBij = TyA,QyB/y = TyA,A/Qy = TyA,-A,Qy = TyA/QyB,y = B/yB,y (S.7) (s>i = 1 k- 1). Since (by definition) A/ is diagonalizable, there exists a nonsingular matrix L; and a diagonal matrix F,- such that Lr1 A,-L/ = F,-, or equivalently such that A/L/ = L/Fi, and hence such that Q-'A/QCQ-'L;) = Q"1 (L/F,) (S.8)
248 21. Eigenvalues and Eigenvectors (/ = 1 k - 1). Comparing (for i = 1 Ic - 1 and j - 1 r) the yth group of rows of the left and right sides of equality (S.8) and using equality (S.6), we find that (0 0, B/y,0 OXT'L, = TyL;F, and hence that Bij(JjL!) = (TjUWi. Thus, any nonnull column of the vy x n matrix TyL,- is an eigenvector of B/y. Moreover, since L, is nonsingular and since the rows of Q"1 are linearly independent, rank(TyL/) = rank(Ty) = vy, so that (according to Theorem 4.4.10) TyL/ contains vy linearly independent columns and it follows from Corollary 21.5.3 that B/y is diagonalizable. We have (for j = 1 r) that each of the matrices Biy Bjt_i.y is diagonalizable, and it follows from result (S.7) that Biy Bjt_i,y commute in pairs. Thus, by supposition, Biy Bjt_i,y can be simultaneously diagonalized; that is, there exists a uy x vj nonsingular matrix Sy and uy x vj diagonal matrices Diy Pft_ity such that (for i = 1 k-l) SJlBUSj=Dtj. (S.9) Define S = diag(Si Sr) and P = QS. Then, clearly, S is nonsingular, and hence P is also nonsingular. Further, using results (S.6), (S.9), and (S.4), we find that,for/ = l /:-1, P-!A|P = S^Q-'A/QS = diag(D/, D,>) and that P"1 AjtP = S-'CT1 AjtQS = diag^I,,, XrlVf), so that all k of the matrices Ai Ajt are simultaneously diagonalized by the nonsingular matrix P. EXERCISE 30. Let V represent an n x n symmetric nonnegative definite matrix, X an n x p matrix of rank r, and d a p x 1 vector. Using the results of Exercise 19.11 (or otherwise), show that each of the following three conditions is necessary and sufficient for the vector X(X;X)_d to be a solution, for every d e C(X'), to the problem of minimizing the quadratic form a'Va (in a) subject to X'a = d: (a) there exists an orthogonal matrix that simultaneously diagonalizes V and Px; (b) there exists a subset of r orthonormal eigenvectors of V that is a basis for C(X); (c) there exists a subset of r eigenvectors of V that is a basis for C(X). Solution, (a) Recall from Part (3) of Theorem 12.3.4 that Px is symmetric. Then, as a consequence of Corollary 21.13.2, there exists an orthogonal matrix that
21. Eigenvalues and Eigenvectors 249 simultaneously diagonalizes V and Px if and only if PXV = VPX. And, it follows from the results of Exercise 19.11 that the existence of an orthogonal matrix that simultaneously diagonalizes V and Px is a necessary and sufficient condition for X(X'X)~d to be a solution [for every d e C(X')] to the problem of minimizing a'Va subject to X'a = d. (b) and (c). Suppose that there exist /* (possibly orthonormal) eigenvectors uj ur of V that form a basis for C(X), and let U = (m ur). Then, VU = UD for some (diagonal) matrix D. Moreover, since clearly C(U) = C(X), X = UT and U = XS for some matrices T and S. Thus, VX = VUT = UDT = XSDT = XQ for Q = SDT. And, it follows from the results of Exercise 19.11 that X(X'X)"d is a solution [for every d e C(X')] to the problem of minimizing a'Va subject to X'a = d. Conversely, suppose that X(X'X)~d is a solution [for every d e C(X')] to the problem of minimizing a'Va subject to X'a = d. Then, it follows from Part (a) that there exists an n x n orthogonal matrix Q that simultaneously diagonalizes Px and V. That is, there exists an n x n orthogonal matrix Q such that Q'PxQ = diag(di d„) and Q'VQ = diag(/i /„) for some scalars d\ d„ and /l /„. Further, it follows from Theorem 21.5.1 that the (not necessarily distinct) eigenvalues of Px are d\ d„ and the (not necessarily distinct) eigenvalues of V are /i /,, and that the first /?th columns of Q are eigenvectors of Px corresponding to d\ dn% respectively, and are also eigenvectors of V corresponding to /i,..., /„, respectively. And, since [according to Part (8) of Theorem 12.3.4] rank(P x) = r, n - r of the eigenvalues of Px are (in light of Lemma 21.1.1) equal toO. Now, let Q! represent the n x r matrix obtained from Q by deleting those columns that are eigenvectors of Px corresponding to 0. Then, in light of Theorem 21.4.3 and Part (7) of Theorem 12.3.4, C(Q,) = C(PX) = C(X). And, since the columns of Qj are (orthonormal) eigenvectors of V, there exist r orthonormal eigenvectors of V that form a basis for C(X). EXERCISE 31. Let A represent an n x n symmetric matrix, and let B represent annxn symmetric positive definite matrix. And, let Xmax and Xmin represent, respectively, the largest and smallest roots of |A - XB|. Show that x'Ax Xmin " x7^ - Xmax for every nonnull vector x in H". Solution. Let S represent any n x n nonsingular matrix such that B = S'S, let R = (S-1)' (so that B"1 = R'R), and let C = RAR;. Then, in light of result (14.7), Xmax and Amjn are, respectively, the largest and smallest eigenvalues of C.
250 21. Eigenvalues and Eigenvectors And, it follows from Theorem 21.5.6 that for every nonull vector y in 1Zn. Now, let x represent an arbitrary nonull vector in Hny and let y = Sx. Then, y is nonull, and y'Cy _ x/S/CSx _ x/Ax y'y ~~ x'S'Sx ~" x'Bx' Thus, x'Ax EXERCISE 32. Let A represent an nxn symmetric matrix, and let B represent an nxn symmetric positive definite matrix. Show that A — B is nonnegative definite if and only if all n (not necessarily distinct) roots of |A - AB| are greater than or equal to 1 and is positive definite if and only if all n roots are (strictly) greater than 1. Solution. Let d\t • • •. d„ represent the n (not necessarily distinct) roots of |A—XB|. And, let S represent any n x n nonsingular matrix such that B = S'S, let R = (S-1)' (so that B"1 = R'R), and let C = RAR'. Then, in light of result (14.7), the (not necessarily distinct) eigenvalues of C are d\ d„, and it follows from Corollary 21.5.9 that there exists an n x n orthogonal matrix P such that P'CP = D, where D = diag(di d„). Now, take Q = S'P. Then, according to results (14.2) and (14.1), A = QD<y and B = QQ;. Thus, A-B = Q(D-IM)Q, = Qdiagtfi-l dn - 1)Q'. And, it follows from Corollary 14.2.15 that A - B is nonnegative definite if and only if, for / = 1 h, d\ - 1 > 0 and is positive definite if and only if, for / = 1 n, d\ — 1 > 0. Or, equivalently, A - B is nonnegative definite if and only if, for i = 1 n, d; > 1 and is positive definite if and only if, for / = 1 H, d; > 1.
22 Linear Transformations EXERCISE 1. Let U represent a subspace of a linear space V, and let S represent a linear transformation from U into a linear space W. Show that there exists a linear transformation T from V into W such that S is the restriction of T to U. Solution. Let {Xi,..., Xr} represent a basis for U. Then, it follows from Theorem 4.3.12 that there exist matrices Xr+i,..., Xr+* such that [X\,..., Xr, Xr+i,..., Xr+it} is a basis for V. Now, fori = 1 /\ define Y,- = 5(X,); and, for / = r + 1 r+k, take Yj to be any matrix in W. And, letting X represent an arbitrary matrix in V, take T to be the transformation from V into W defined by T(X) = cj Yj + • • • + crYr + Cr+\ Yr+, +...+ cr+kYr+k, where c\ cr, cr+\,..., cr+k are the (unique) scalars that satisfy X = qXj + • ■ • + crXr + cr+jXr+i +•••+ cr+JtXr+jt — since Yj Yr, Yr+l Yr+k are in the linear space W, T(X) is in W. Clearly, if X e U, then cr+\ = • • - = cr+k = 0, and hence T(X) = c,Y, +.-. + crYr = f^CiS(Xi) = slj^aXi) = S(X). Moreover, it follows from Lemma 22.1.8 that T is linear. Thus, T is a linear transformation from V into W such that S is the restriction of T to U. EXERCISE 2. Let T represent a 1-1 linear transformation from a linear space V into a linear space W. And, write U* Y for the inner product of arbitrary matrices
252 22. Linear Transformations U and Y in W. Further, define X * Z = T(X) • T(Z) for all matrices X and Z in V. Show that the "*-operation" satisfies the four properties required of an inner product for V. Solution. Observe (in light of Lemma 22.1.3) that J\f(T) = {0} and hence that T(X) = 0 if and only if X = 0. Then, letting X, Z, and Y represent arbitrary matrices in V and letting k represent an arbitrary scalar, we find that (1) X*Z = 7(X)-7(Z) = 7,(Z)-7(X) = Z*X; (2) X*X = 7,(X)-7'(X)>0, if7(X)#0or,equivalently,ifX#0, = 0, if T(X) = 0 or, equivalently, if X = 0; (3) (kX)*Z = T{kX)*T(Z) = [kT(X)]*T(Z) = k[T(X)*T(Z)] = k(X* Z); and (4) (X + Z)*Y = r(X + Z)-7(Y) = [7(X) + 7(Z)]-r(Y) = [nX)-T(Y)) + [nZ)-:T(Y)] = (X*Y) + (Z*Y). EXERCISE 3. Let T represent a linear transformation from a linear space V into a linear space W, and let U represent any subspace of V such that U and N{T) are essentially disjoint. Further, let {Xj Xr} represent a linearly independent set of /-matrices in U. (a) Show that T(X\) T(Xr) are linearly independent. (b) Show that if r = dim(W) (or equivalently if Xj Xr form a basis forU) and if U 0 N{T) = V, then T(X\) T(Xr) form a basis for T(V). Solution, (a) Let ci cr represent any scalars such that ^=1 c,T(X,-) = 0. Then, n£/=1 c«X,-) = £J=1 c,T(X/) = 0, implying that £[=1 cfX/ is in M(T) and hence (since clearly ^=1 c,X,- e U) that ^=1 c,X/ is in UnJ\f(T). Thus, 53[=j qXj = 0, and (since Xj Xr are linearly independent) it follows that C] =-.. = 0 = 0. And, we conclude that ^(Xi) T(Xr) are linearly independent. (b) Suppose that /• = dim(W) and that U®NiT) = V. Then, making use of Theorem 22.1.1 and of Corollary 17.1.6, we find that dimfnV)] = dim(V) - dim^n] = r. And, in light of Theorem 4.3.9 and the result of Part (a), it follows that ^(Xi), ..., r(Xr) form a basis for T{V). An alternative proof [of the result of Part (b)] can be obtained [in light of the result of Part (a)] by showing that the the set {^(Xi) T(Xr)} spans T(V). Continue to suppose that r = d\m(U) and lhatU (&Af(T) = V. And, let Zj Zs represent any matrices that form a basis for N(T). Further, observe, in light of Theorem 17.1.5, that the r + s matrices Xj Xr, Zj Zs form a basis for V. Now, let Y represent an arbitrary matrix in ^V). Then, Y = ^(X) for some
22. Linear Transformations 253 matrix X in V and X = ££_, qX/ + £}=1 kjzj>so that Y = MEc* + E*7Z; ) = X/'^X,) + E^r^> = X>n*>. \i=l y=I / /=1 y=I »=l Thus, {^(Xi) T(Xr)} spans T(V). EXERCISE 4. Let 7 and 5 represent linear transformations from a linear space V into a linear space W, and let k represent an arbitrary scalar. (a) Verify that the transformation kT is linear. (b) Verify that the transformation T 4- S is linear. Solution, (a) For any matrices X and Z in V and for any scalar c, (kT)(X + Z) = kT(X + Z) = T[k{X + Z)] = T(kX + kZ) = T{kX) + T(kZ) = kT(X) + kT(Z) = (kT)(X) + (kT)(Z). and (kT)(cX) = kT(cX) = k[cT(X)] = c[kT(X)] = c(kT)(X). (b) For any matrices X and Z in V and for any scalar c\ (T + S)(X + Z) = T(X + Z) + S(X + Z) = r(X) + r(Z) + S(X) + S(Z) = r(X) + 5(X) + r(Z) + 5(Z) = (7 + S)(X) + (7 + SMZ), and (7 + 5)(cX) = T(cX) + 5(cX) = cT(X) + cS(X) = c[7(X) + S(X)] = c(T + S)(X). EXERCISE 5. Let S represent a linear transformation from a linear space U into a linear space V, and let T represent a linear transformation from V into a linear space W. Show that the transformation TS is linear. Solution. For any matrices X and Z in U and for any scalar c, (75)(X + Z) = 7[5(X + Z)] = T[S(X) + 5(Z)] = 7[5(X)] + T[S(Z)] = (75)(X) + (7\S)(Z), and (75)(cX) = 7[5(cX)] = T[cS(X)] = c7[5(X)] = c(TS)(X).
254 22. Linear Transformations EXERCISE 6. Let T represent a linear transformation from a linear space V into a linear space W, and let R represent a linear transformation from a linear space U into W. Show that if T(V) C R(U\ then there exists a linear transformation S from V into U such that 7 = RS. Solution. Suppose that T(V) C R{U). And, let {Xi Xr} represent a basis for V. Then, for / = 1,..., r, HX,-) e #(£/), and consequently 7(¾) = tf(Y,) for some matrix Y/ mli. Now, let X represent an arbitrary matrix in V, and let c\,..., cr represent the (unique) scalars that satisfy the equality X = £/=i c,X,-. And, take S to be the transformation from V into U defined by S(X) = £JL, c,-Y/. Then, T(X) = T[ £c,X, ) = £^7-(¾) \/=i / /=1 r = Y,ciR(Vi) i=\ = ^(Ec'Y') = RV{X)] = (^xx)- Thus, T = jR5. Moreover, it follows from Lemma 22.1.8 that S is linear. EXERCISE 7. Let T represent a transformation from a linear space V into a linear space W, and let S and 7? represent transformations from W into V. And, suppose that RT = I (where the identity transformation I is from V onto V) and that TS = I (where the identity transformation I is from W onto W). (a) Show that T is invertible. (b) Show that R = S = T~\ Solution, (a) For any matrices X and Z in V such that T(X) = T{Z), X = /(X) = (RT)(X) = R[T(X)] = R[T(Z)] = {RT)(Z) = /(Z) = Z. Thus, T is 1-1. Further, for any matrix Y in W, Y = I(Y) = (TS)(Y) = T(X), where X = S(Y). And, it follows that T is onto. Since T is both 1-1 and onto, we conclude that T is invertible. (b) Using results (3.3) and (3.1), we find that R = RI = R{TT~l) = (RT)T~l = IT'1 = T~l and S = IS = (T~lT)S = T~l(TS) = T~lI = T~\
22. Linear Transformations 255 EXERCISE 8. Let T represent an invertible transformation from a linear space V into a linear space W, let 5 represent an invertible transformation from a linear space U into V, and let k represent an arbitrary scalar. Using the results of Exercise 7 (or otherwise), show that (a) kT is invertible and (AT)-1 = (1//:)7-1 and that (b) TS is invertible and (TS)~l = S~lT-\ Solution, (a) In light of the results of Exercise 7, it suffices to show that ((1//:)7^)(/:7-) = / and {kT){{\/k)T~x) = I. Using results (2.12), (3.1), (3.3), (2.2), and (2.1), we find that ((l/k)T-l)(kT) = (l/k)(T-l(kT)) = (l/k)(k(T-lT)) = (l/k)(kl) = [(\/k)k]I = 1/ = / and similarly that (kT)«l/k)T-1) = (l/k)((kT)T-1) = (l/k)(k(TT-1)) = (l/k)(kl) = [(1/*)*]/ = 1/ = /. (b) In light of the results of Exercise 7, it suffices to show that (S-lT~l)(TS) = I and (TS)(S-lT"l) = I. Using results (2.9), (3.1), (3.3), and (2.13), we find that (S~lT'l)(TS) = aS~lT-l)T)S = (S-l(T~lT))S = (S~lI)S = S~lS = I and similarly that (75)(5-^-1) = ((TS)S~l)T-1 = (7(55-1))^1 = (77)7-1 = 7-7--1 = /. EXERCISE 9. Let T represent a linear transformation from an H-dimensional linear space V into an w-dimensional linear space W. And, write U • Y for the inner product of arbitrary matrices U and Y in W. Further, let B represent a set of matrices Vi,..., V„ (in V) that form a basis for V, and let C represent a set of matrices Wi Wm (in W) that form an orthonormal basis for W. Show that the matrix representation of T with respect to B and C is the mxn matrix whose ijth element is T(Yj)*Wi. Solution. As a consequence of Theorem 6.4.4, we have that (for j = 1,..., n) m
256 22. Linear Transformations And, upon comparing this expression for T(\j) with expression (4.3), we find that the matrix representation of T with respect to B and C is the m x n matrix whose ijth element is T{Vj )• W,-. EXERCISE 10. Let T represent the linear transformation from W'x" into U" xm defined by T(X) = X'. And, let C represent the natural basis for ft,"'*", comprising the mn matrices V\ \, U21 UW| Ui„, U2,, U„„,. where (for / = 1 /» and j = 1,... n) U,y- is the m x n matrix whose ijth element equals 1 and whose remaining mn - 1 elements equal 0; and similarly let D represent the natural basis for TZnx'". Show that the matrix representation for T with respect to the bases C and D is the vec-permutation matrix K„„,. Solution. Making use of results (4.11) and (16.3.1), we find that, for any m x n matrix X, (L^FLcKvec X) = vec[r(X)] = vec(X') = KHI„vec X. And, in light of result (4.7), it follows that the matrix representation of T (with respect to C and D) equals Kmn . EXERCISE 11. Let W represent the linear space of all pxp symmetric matrices, and let T represent a linear transformation from 11'"*" into W. Further, let B represent the natural basis for TZmx"% comprising the mn matrices Un, U21 U„,i Ui„, U21, U„,„, where (for i = 1 /;* and j = 1 n) U,y is the in x n matrix whose ijth element equals 1 and whose remaining mn - 1 elements equal 0. And, let C represent the usual basis for W. (a) Show that, for any m x n matrix X, (LclTLB)(vec X) = vech[r(X)]. (b) Show that the matrix representation of T (with respect to B and C) equals the p(p + 1)/2 x mn matrix [vech T(VU) vech 7(11,,,1) vech ^(Ui,,) vech 7(11,,,,,)]. (c) Suppose that p = m = n and that (for every n x n matrix X) 7^) = (1/2)^ + ^). Show that the matrix representation of T (with respect to B and C) equals (GiCr'G; (where G„ is the duplication matrix). Solution, (a) Making use of result (3.5). we find that, for any mn x 1 vector x, (L^TLB)(\) = (L~l(TLB))(\) = Lcl[iTLB)(x)\ = vech[(7,L/?)(x)] = vech{r[Lfi(x)]}.
22. Linear Transformations 257 And, in light of result (3.4), it follows that, for any m x n matrix X, (L^lTLB)(vecX) = vech{7,(L/?(vec X)l} = vech(T{LB[L^l(X)]}) = vecn^X)]. (b) For any /» x n matrix X = {.y,,}, we find [using Part (a)] that (L^TLbUvcc X) = vech 71 JjJCf/Uy ) = vech £^70¾) = £\v,yvech[T(U,v)] 'J = [vech T(UU) vech T(V,„\), ..., vech 7(Ui„) vech 7,(Ul„„)]vec(X). And, in light of result (4.7), it follows that the matrix representation of T (with respect to B and C) equals the /7(/7 + 1 )/2 x mn matrix [vech HUii) vech T(VM\) vech T{V\„) vech TiVmn)]. (c) Using Part (a) and results (16.4.6), (16.3.1), and (16.4.15), we find that, for any n x n matrix X, (LclTLB)(vzc X) = vech[(l/2)(X + X')] = (l/2)vech(X + X') = (l/2)(Gl,,Gl,)-1Gl,,vec(X + X') = (l^KG^Cr'G^vecfX) + K„„vec(X)] = (l^HfG^Cr'G,', + (G^Gj-'G^K^Kvec X) = (l/2)[{G'„G„r]G'n + (GiG^r'GiKvec X) = (Gl,,Gl,)-,Gl,,(vecX). And, in light of result (4.7), it follows that the matrix representation of T (with respect to B and C) equals (G^G,,)-^,. EXERCISE 12. Let V represent an /1-dimensional linear space. Further, let B = {Vi, V2 V,,} represent a basis for V, and let A represent an n x n nonsingular matrix. And, for j = 1 /1. let Wy = fij\i + /2;V2 + ••■ + f„jV„ , where (for i = 1 n) fij is the ijih element of A-1. Show that the set C comprising the matrices Wj, Wo,... W„ is a basis for V and that A is the matrix
258 22. Linear Transformations representation of the identity transformation / (from V onto V) with respect to B andC. Solution. Lemma 3.2.4 implies that the set C is linearly independent and hence that C is a basis for V. Then, clearly, A-1 is the matrix representation of the identity transformation I (from V onto V) with respect to C and B. And, it follows from Corollary 22.4.3 that (A-1)-1 is the matrix representation of I~l with respect to B and C and hence [since (A-1)-1 = A and I~l = I] that A is the matrix representation of I with respect to B and C. EXERCISE 13. Let T represent the linear transformation from H4x l into V?x l defined by T(\) = (X[ + A"2, x2 4- *3 - *4» *i - *3 + *4>'. where x = Ui, *2»*3. *4)'- Further, let B represent the natural basis for TZ4xl (comprising the columns of I4), and let E represent the basis (for Tl4x l) comprising the four vectors (1, -1.0, -1)\ (0,0,1,1)', (0,0,0, 1)', and (1.1,0,0)'. And, let C represent the natural basis for V?x l (comprising the columns of I3), and F represent the basis (for TZ3xl) comprising the three vectors (1,0,1)', (1,1,0)', and (-1,0,0)'. (a) Find the matrix representation of T with respect to B and C. (b) Find (1) the matrix representation of the identity transformation from TZ4xl onto 1Z4x l with respect to E and B and (2) the matrix representation of the identity transformation from 1Zixl onto V?x l with respect to C and F. (c) Find the matrix representation of T with respect to E and F via each of two approaches: (1) a direct approach, using the equality WA = [7(V,) T(V„)l (*) where A is the matrix of a linear transformation T from an ^-dimensional linear space V into 11'" x l, where {Vi V„} is the basis for V, and where W is an »1 x m matrix whose columns form the basis for TV" x l; and (2) an indirect approach, using the results of Parts (a) and (b) in combination with the result that if A is the matrix representation of a linear transformation T from a linear space V into a linear space W with respect to bases B and C (for V and W, respectively), then the matrix representation of T with respect to alternative bases E and F is S_1AR, where R is the matrix representation of the identity transformation from V onto V with respect to E and B and S is the matrix representation of the identity transformation from W onto W with respect to F and C. (d) Find rank T and dimLVCT)]. Do so by, for instance, using the result that the rank of a linear transformation T from an /1-dimensional linear space V into a linear space W equals the rank of its matrix representation (with respect to any bases B and C) and the result that dim[^(7)1 =//- ranker).
22. Linear Transformations 259 Solution, (a) Let A represent the matrix representation of T with respect to B and C, and denote the first 4th columns of I4 by ei,..., e4, respectively. Then, in light of the discussion in Part 1 of Section 4b, we find that A = I3A = [r(ei) 7Xe4)] (1 1 0 0\ 0 1 1 -1 . 10-1 1/ (b) (1) In light of the discussion in Part 3 of Section 4b, the matrix representation of the identity transformation from TZ4xl onto TZ4xl with respect to E and B is the 4 x 4 matrix whose first 4th columns are the vectors that form E, that is, the 4 x 4 matrix ( 1 -1 0 l-l 0 0 1 1 0 0 0 1 1 0 0 (2) Let S represent the 3 x 3 nonsingular matrix whose inverse S_1 is the matrix representation of the identity transformation from TZ3xl onto fc3xl with respect to C and F. Then, in light of Part 3 of Section 4b, so that /1 1 -a 0 1 0 s-'=i3 / 0 0 i\ s-'= 0 1 0 . V-i 1 1) (c) Let H represent the matrix representation of T with respect to E and F. (1) Equality (*) [or equivalently equality (4.8)] gives (1 1 -l\ /0 0 0 2\ 0 1 0 H= 0 0-11. 10 0/ \0 0 1 1/ Thus, (00 1 1\ 0 0-11. 0 0 0 0/ (2) The matrix 5 [from Part (b)] is (in light of Corollary 22.4.3) the matrix representation of the identity transformation from TZ3xl onto 7£3xl with respect to F and C. Thus, in light of the results of Parts (a) and (b), it follows from the
260 22. Linear Transformations result cited (which is Theorem 22.4.4) that -(J ::)(11-:¾ (00 1 1\ 0 0-1 1 . 0 0 0 0/ OOP 0 0 1 1 0 0 1 1 oj (d) It is clear from Part (c) that the rank of the matrix representation of T with respect to E and F equals 2. Thus, it follows from the first result cited (which is Theorem 22.5.2) that rank T = 2 and from the second result cited (which is part of Corollary 22.5.3) that dim [Af(T)) = 4 - 2 = 2. EXERCISE 14. Let T represent a linear transformation of rank k (where k > 0) from an n-dimensional linear space V into an m-dimensional linear space W. Show that there exists a basis E for V and a basis F for W such that the matrix representation of T with respect to E and F is of the form ( * ft J. Solution. Let B represent any basis for V and C any basis for W. And, let A represent the matrix representation of T with respect to B and C. Then, as a consequence of Theorem 22.5.2, rank A = k, and it follows from Theorem 4.4.9 that there exists an n x n nonsingular matrix R and an m x in nonsingular matrix S such that A = S ( * 0 ) R"l or equivalently such that ( * Q J = S~l AR. And, based on Theorem 22.4.7, we conclude that I ' ft J is the matrix representation of T with respect to some bases E and F. EXERCISE 15. Let T represent a linear transformation from an ^-dimensional linear space V into an m-dimensional linear space W, and let A represent the matrix representation of T with respect to bases B and C (for V and VV, respectively). Use the result that an n x 1 vector x is in MA) if and only if the corresponding matrix Lg(x) is in J\f(T) [or equivalently that a matrix X (in V) is in J\f{T) if and only if the corresponding vector L^!(X) is in MA)] to devise a "direct" proof that dim[Af{T)] = dimLVfA)] (as opposed to deriving this equality as a corollary of the result that rank T = rank A). Solution. The result cited [which is Part (2) of Theorem 22.5.11 implies that LflLV(A)] = Af(T){2iS can be easily verified). Thus, there exists a 1-1 linear transformation fromM{A) ontoN(T), namely, the linear transformation R defined [for x e Af(A)] by R(\) = LB(\). And, it follows that M{A) and N(T) are isomorphic. Based on Theorem 22.3.1, we conclude that dim[.A/'(n] = dim[A"(A)].
22. Linear Transformations 261 EXERCISE 16. Let T represent a linear transformation from an n-dimensional linear space V into V. (a) Let U represent an r-dimensional subspace of V, and suppose that U is invariant relative to T. Show that there exists a basis B for V such that the matrix representation of T with respect to B and B is of the (upper block-triangular) form (E F\ ft „ J (where E is of dimensions r x r). (b) Let U and W represent subspaces of V such that U © W = V (i.e., essentially disjoint subspaces of V whose sum is V). Suppose that both U and W are invariant relative to T. Show that there exists a basis B for V such that the matrix representation of T with respect to B and B is of the (block-diagonal) form diag(E, H) [where the dimensions of E equal dim(VV)]. Solution, (a) Let Xi Xr represent any r matrices that form a basis for U. And, take B to be any basis for V comprising Xj Xr and n — r additional matrices Xr+i X„ — the existence of such a basis is guaranteed by Theorem 4.3.12. Further, let A = {atj} represent the matrix representation of T with respect to B and B. Then, for j = 1 r, a\jXi H h orjXr + flr+i.yXr+i H h a„jX„ = T(Xj) e U, implying (since any matrix in li can be expressed as a linear combination of Xi Xr) that tfr+l j = " = anj = 0. Thus, a/y = 0 for i = r + 1 n and 7 = 1 r. (b) Let r = dim(W) [in which case dim(W) = n — r]. Further, let Xi X,. represent any r matrices that form a basis for U and Xr+i X„ any n—r matrices that form a basis for W. And, take B to be the basis for V comprising Xi,..., Xr, Xr+i X„ — that Xi, ..., Xr, Xr+i X„ form a basis for V is evident from Theorem 17.1.5. Now, let A = {cijj} represent the matrix representation of T with respect to B and B. Then, for j = 1 /\ r + 1,..., n, a\jX\ H 1-arjXr + fl,+i,;Xr+i H ha„jX„ = T(Xj). And, for j = 1 r, HXy) e U, implying (since any matrix in U can be expressed as a linear combination of Xi X,) that ar+ij = • • • = anj = 0. Similarly, for j = r+1 n, T(Xj) e W, implying (since any matrix in Wean be expressed as a linear combination of Xr+i,..., X„) that a\j = • • • = arj = 0. Thus, a-,j = 0 for / = r + 1,..., n and j = 1 r, and also a\j = 0 for i = 1,..., r and j = r + 1,..., n. EXERCISE 17. Let V, W, and U represent linear spaces. (a) Show that the dual transformation of the identity transformation I from V onto V is /. (b) Show that the dual transformation of the zero transformation 0j from Vinto W is the zero transformation O2 from W into V.
262 22. Linear Transformations (c) Let 5 represent the dual transformation of a linear transformation T from V into W. Show that T is the dual transformation of 5. (d) Let k represent a scalar, and let 5 represent the dual transformation of a linear transformation T from V into W. Show that kS is the dual transformation of*7\ (e) Let T\ and 7¼ represent linear transformations from V into W, and let S\ and .% represent the dual transformations of T\ and 72, respectively. Show that S\ 4- 52 is the dual transformation of T\ 4- 7½. (f) Let P represent the dual transformation of a linear transformation 5 from U into V, and let Q represent the dual transformation of a linear transformation T from V into W. Show that PQ is the dual transformation of TS. Solution. Write X • Z for the inner product of arbitrary matrices X and Z in V, U * Y for the inner product of arbitrary matrices U and Y in W, and A * B for the inner product of arbitrary matrices A and B in U. (a) For every matrix X in V and every matrix Y in V, X-/(Y) = X-Y = /(X)-Y. Thus, / is the dual transformation of /. (b) For every matrix X in V and every matrix Y in W, X-02(Y) = X-0 = 0 = 0 * Y = Oi(X) * Y. Thus, O2 is the dual transformation of Oi. (c) For every matrix Y in W and every matrix X in V, Y * T(X) = T(X) * Y = X-S(Y) = S(Y)-X. Thus, T is the dual transformation of S. (d) For every matrix X in V and every matrix Y in W, X-(kS)(Y) = X-[kS(Y)] = *[X-S(Y)] = k[T(X) * Y] = [kT(X)] * Y = (kT)(X) * Y, Thus, kS is the dual transformation of kT. (e) For every matrix X in V and every matrix Y in W, X-(S, +S2)<Y) = X-[Si(Y) + S2(Y)] = X-S|(Y) + X-S2(Y) = 7,,(X)*Y+72(X)*Y = ffi(X) + T2(X)1 * Y = (7, + 72)(X) * Y. Thus, S\ + S2 is the dual transformation of T\+T2.
22. Linear Transformations 263 (f) For every matrix X in U and every matrix Y in W, X* (P0(Y) = X * P[Q(Y)] = S(X)-Q(Y) = T[S(X)] * Y = (TS)(X) * Y. Thus, PQ is the dual transformation of TS. EXERCISE 18. Let S represent the dual transformation of a linear transformation T from an n-dimensional linear space V into an w-dimensional linear space W. And, let A = {a,j} represent the matrix representation of T with respect to orthonormal bases C and D, and B = {£,/} represent the matrix representation of S with respect to D and C. Using the result of Exercise 9 (or otherwise), show that B = A'. Solution. Write X*Z for the inner product of arbitrary matrices X and Z in V, and U * Y for the inner product of arbitrary matrices U and Y in W. And, let Xi X„ represent the matrices that form the orthonormal basis C and Yi Ym the matrices that form the orthonormal basis D. Then, for / = 1, ..., m and ./ = 1, ..., /i, it follows from the result of Exercise 9 that aij = T(Xj)*Yi and bji = SCtfrXj, implying that bji = Xj .5(Y,) = T(Xj) * Y/ = aij and hence that the /7 th element of B equals the /7th element of A'. Thus, B = A'. EXERCISE 19. Let A represent anmxn matrix, let V represent an n xn symmetric positive definite matrix, and let W represent anmxm symmetric positive definite matrix. And let S represent the dual transformation of the linear transformation T from ft" xl into TZmxl defined by T(x) = Ax (where x is an arbitrary/7 x 1 vector). Taking the inner product of arbitrary vectors x and z in TZ"xl to be x'Vz and the inner product of arbitrary vectors u and y in TZmxl to be u'Wy, obtain a formula for S(y) that generalizes the formula S(y) = A'y (y e TZmxl) obtained in the special case of the usual inner products (i.e., in the special case where V = l„ and W = Im). Solution. For every vector x in TZnxl and every vector y in TZmx\ x'\S(y) = (Ax)'Wy = x'A'Wy. And, in light of the uniqueness of 5, 5(y)=V"1A,Wy. EXERCISE 20. Let S represent the dual transformation of a linear transformation T from a linear space V into a linear space W.
264 22. Linear Transformations (a) Show that [S(W)]-1- = N{T) (i.e., that the orthogonal complement of the range space of S equals the null space of T). (b) Using the result of Part (c) of Exercise 17 (or otherwise), show that [M(S)]1- = T(V) (i.e., that the orthogonal complement of the null space of 5 equals the range space of T). (c) Show that rank S = rank T. Solution, (a) Write X'Z for the inner product of arbitrary matrices X and Z in V and U * Y for the inner product of arbitrary matrices U and Y in W. Let X represent an arbitrary matrix in V. Suppose that X € N(T). Then, for every matrix Y in W, X-S(Y) = T(X) *Y = 0*Y = 0. Thus, X 6 [5(W)]X. Conversely, suppose that X e [SQ/V)]1. Then, since T{X) e W, T(X) * T(X) = X-StnX)] = 0, implying that T(X) = 0 and hence that X € N{T). Thus, [SiW)]1 =N{T). (b) Since [according to Part (c) of Exercise 17] T is the dual transformation of 5, it follows from Part (a) that [TCV))1 = J\f(S). And, making use of Theorem 12.5.4, we find that [MS)]1 = [[TiV)]1)1 = T(V). (c) Making use of Corollary 22.5.3 and Theorem 12.5.12 [together with the result of Part (a)], we find that rank T = dim(V) - dim[Af{T)] = dim(V)-dim{[S(W)]±} = dim(V) - (dim(V) - dim[S(W)]} = dim[S(W)l = rank S.
References References Bartle, R. G. (1976), The Elements of Real Analysis (2nd ed.). New York: John Wiley. Goodnight, J. H. (1979), "A Tutorial on the SWEEP Operator," The American Statistician, 33, 149-158. Magnus, J. R., and Neudecker, H. (1980), 'The Elimination Matrix: Some Lemmas and Applications," SIAM Journal on Algebra and Discrete Mathematics, 1,422-449. Meyer, CD. (1973). "Generalized Inverses and Ranks of Block Matrices," SIAM Journal on Applied Mathematics, 25, 597-602.
Index adjoint matrix, 72, 75, 232 determinant of, see under determinant differentiation of, see under differentiation eigenvalues of, see under eigenvalue(s) eigenvectors of, see under eigenvectors) of a product, 77 algebraic multiplicity. 236 of zero, 235,236 basis, 15 orthonormal,22,24,31 bilinear form, 79 Binet-Cauchy formula, 77 Cayley-Hamilton theorem, 232,234 characteristic polynomial, 232,233,243 of an orthogonal matrix, 242 cofactor matrix, 72 cofactor(s), 74 expansion by, see under determinant column space(s), 13,15,31,55,167 essential disjointness of, see under essential disjointness intersection of, 161 of a product, 27 of a sum, 198 orthogonal complement of, see under orthogonal complement sum of, 161 union of, 161 decomposition Cholesky, 90,91 LDU, 87-91,98,144 of a nonnegative definite matrix, 86 of a symmetric matrix, 30,83 QR, 24,90,91 singular value, 245,246 spectral, 239, 241 U'DU, 87 determinant, 71,87 differentiation of, see under differentiation effect of elementary row or column operations on, 70 expansion by cofactors, 71,74 of a partitioned matrix, 71,76 of a positive definite matrix, 100 of a product, see Binet-Cauchy formula of an adjoint matrix, 72
268 Index of an inverse matrix, 71 ofR + STU,179,180 of Vandermonde matrix, 77 diagonalization, 238 of a transposed matrix, 237 of an inverse matrix, 237 simultaneous, 246 differentiation chain rule for, 122-124 of a determinant, 124 of a Kronecker product, 158 of a log of a determinant, 125-128, 131,132 of a power of a determinant, 124 of a power of a function, 115 of a product of a scalar and a vector, 116 of a projection matrix, 135,136 of a trace of a power, 119,120 of a trace of a product, 116, 117, 120 of a vec of a Kronecker product, 158 of a vec of a matrix power, 157 of an adjoint matrix, 129 of an idempotent matrix. 115 of an inverse matrix, 130-132 with respect to a matrix or symmetric matrix, 122,125,126, 128,130-132 distance. 22 duplication matrix, 154-156 left inverse of, 154-156 eigenvalue(s), 231,243 algebraic multiplicity of, see algebraic multiplicity geometric multiplicity of, see geometric multiplicity not necessarily distinct, 238,240 of a skew-symmetric matrix, 231 of an adjoint matrix, 242 of an idempotent matrix, 242 of an orthogonal matrix, 242 of Moore-Penrose inverse, 241 eigenvectors), 243 linear combination of, 237 of an adjoint matrix, 242 elimination matrix, 154 essential disjointness (of subspaces), 252 as applied to row and column spaces, 168,198 function (continuously differentiable), 113 generalized eigenvalue problem, 249,250 generalized inverse, 35,36,91,193,200 alternative characterizations for, 35, 37 existence of, 35 minimum norm, 223 nonnegative definite, 87 of a block-diagonal matrix, 39,45 of a partitioned matrix, 39,41,42, 45,46,52,97,169,170,222 of a product, 51,106 of a scalar multiple, 38 of a Schur complement, 46 of a submatrix, 45,46 of A'A, 36 ofR + STU, 185,186 reflexive, 222 geometric multiplicity of one, 242 of zero, 236 Gram matrix, 81 Gram-Schmidt orthogonalization, 22-24. 66 Gramian, 81 Hadamard product, 95 Hessian matrix, 114 index of inertia, 83 inner product, 102,105, 147,251 quasi, 105 inverse, 29,30,74, 234 determinant of, see under determinant diagonalization of, see under diagonalization differentiation of, see under differentiation of a 2 x 2 matrix, 73 of a block-triangular matrix, 31 of a positive definite matrix, 82,182 of a sum or difference, 184, 187, 188, 191
Index 269 ofR + STU, 184 Kronecker product, 140,156 differentiation of, see under differentiation generalized inverse of, 141 involving a diagonal matrix, 156 involving a partitioned matrix, 143 involving a sum (or sums), 139 involving a triangular matrix, 144, 156 involving a vector, 140,151, 154 LDU decomposition of, 144 nonnegative definite, 142 norm of, 143 of idempotent matrices, 140 of orthogonal matrices, 140 of projection matrices, 141 positive definite, 142 projection matrix for column space of, see under projection matrix vec of, see under vec left inverse* 29 of duplication matrix, see under duplication matrix linear dependence, 11,12 linear independence. 11.12.145 linear space(s), 13 basis for, see basis isomorphic, 260 subspace(s) of, see under subspace(s) linear system(s) absorption in. 211 augmented. 60 consistent. 58 Cramer's rule for, 75 equivalence of, 58,157,210 homogeneous, 55 inconsistent, 58 invariance to choice of solution of, 60 linear combination of solutions of. 55,56 nonhomogeneous, 56 oftheformX'XB = X',65 solution set of, 56, 58 solution to, 57,75 linear transformation(s). 252 dual, 261, 263 identity, 258, 261 matrix representation of, 255,256, 258,260,261,263 null space of, 260,264 one to one, 251 product of, 253,254,262 range space of, 264 rank of, 264 restriction of, 251 scalar multiple of, 253,262 sum of, 253, 262 zero, 261 matrix (or matrices) commutativity of, 3 congruence of, 83 diagonal, 7 diagonally dominant, 98 difference between, 3 idempotent, 49, 50, 82, 115, 140, 146,189-192,195,226 invertible, 29 involutory, 29, 50 negative definite, 80,83,142 negative semidefinite, 80 nonnegative definite, 84,90,93,102, 142.180,181,188-190,195, 198,250 2x2, 102 partitioned, 96,97 sum of, 93 nonpositive definite, 80, 142 nonsingular, 30,98,135 nonsymmetric, 13 of the form V + XUX', 209,212 orthogonal, 30, 31, 49, 140, 146, 221 permutation, 149 positive definite, 80,83,84,98,101, 105,135, 142, 191, 250 2x2, 101 product of, 94 positive semidefinite, 80,90 nonsingular, 80 partitioned, 96 power of, 4 product of, 1-3, 27
270 Index scalar multiple of, 1 similarity of, see similarity singular, 72 skew-symmetric, 92,93,183 submatrix of, see submatrix sum of, 1-3 symmetric, 3,7,20,36,72 transpose of, 4,9 triangular, 31 lower, 4, 144 upper, 4,7,8,13,79,144 minimization (of a 2nd-degree polynomial), 209 subject to linear constraints, 214, 216-218,225,248 transformation from constrained to unconstrained, 214, 227 Moore-Penrose conditions, 223 Moore-Penrose inverse, 221,223,225 eigenvalues of, see under eigenvalue(s) of a nonnegative definite matrix, 228 of a product, 221,225 of a sum, 224 of a symmetric nonnegative definite matrix, 226 neighborhood, 113 norm, 21 limit of, 187 quasi, 105 usual. 143 normal equations, 65 null space. 112 orthogonal complement, 177 dimension of, 112 of a column space, 112 of a sum, 162 of an intersection, 162 projection on, 67,112 orthogonality of 2 subspaces, 63, 162,177 of a matrix and a subspace, 63,64, 162 of a vector and a subspace, 107 partitioned matrix (or matrices) block-triangular. 8 determinant of, see under determinant generalized inverse of, see under generalized inverse in a Kronecker product, see under Kronecker product nonnegative definite, see under matrix (or matrices): nonnegative definite positive semidefinite, see under matrix (or matrices): positive semidef- inite product of, 9 rank of, see under rank row space of, see under row space(s) Schur complement in, see Schur complement transpose of, 9 positive or negative pair (of matrix elements), 69 projection, 64 along a subspace, 173,177 of a column vector, 65,106,111 on an orthogonal complement, see w/N&rorthogonal complement projection matrix, 66,108-110,213 differentiation of, see under differentiation for column space of a Kronecker product, 141 for one subspace along another, 173- 175, 193, 212 in a Kronecker product, see under Kronecker product quadratic form (matrix of), 79 rank, 15,17,82 additivity of, 191, 192, 195, 199, 203 full column. 30 full row, 30 of a difference, 200 of a partitioned matrix, 17,51,53, 97, 167, 168,171, 204 of a product, 16,27,172 of a sum, 197,198 of a triangular matrix, 31 of R + STU, 196,206
Index 271 subtractivity of, 200 right inverse, 29, 174 row space(s), 13 of a partitioned matrix, 17 of a product, 27 of a sum, 198 Schur complement, 42,46,98 generalized inverse of, see under generalized inverse Schwarz inequality, 21 set interior point of, 114 open, 113,135 span of, 14,15,163 similarity, 234,235 to an idempotent matrix, 234 submatrix generalized inverse of, see under generalized inverse principal, 7,98 transpose of, 7 subspace(s), 14 direct sum of, 261 essential disjointness of, see essential disjointness independence of, 164,173,177,203 intersection of, 163 invariant, 231,261 orthogonality of, see orthogonality projection along, see under projection projection matrix for, see underprojection matrix sum of, 161-164, 170 union of, 161 sweep operation, 43 differentiation of, see under differentiation of a product, 19,94,146 of a sum, 96 transformation(s) inverse of, 254, 255 invertible, 254, 255 linear, see linear transformation(s) product of, 255 scalar multiple of, 255 triangle inequality, 21,22,187 vec, 114, 155, 157 differentiation of, see under differentiation of a Kronecker product, 153 of an idempotent matrix, 146 of an identity matrix, 145 of an orthogonal matrix, 146 vec-permutation matrix, 151-153,256 determinant of, 149 recursive formula for, 149 vech, 154,155,157
This book contains over 300 exercises and solutions covering a wide variety of topics in matrix algebra. They can be used for independent study or in creating a challenging and stimulating environment that encourages active engagement in the learning process. Thus, the book can be of value to both students and teachers. The requisite background is some previous exposure to matrix algebra of the kind obtained in a first course. The exercises are those from an earlier book by the same author entitled Matrix Algebra From a Statistician's Perspective. They have been restated as necessary to stand alone, and the book includes extensive and detailed summaries of all relevant terminology and notation. The coverage includes topics of special interest and relevance in statistics and related disciplines, as well as standard topics. The overlap with exercises available from other sources is relatively small. David A. Harville is a research staff member in the Mathematical Sciences Department of the IBM T.J. Watson Research Center. Prior to joining the Research Center, he served ten years as a mathematical statistician in the Applied Mathematics Research Laboratory of the Aerospace Research Laboratories at Wright-Patterson Air Force Base, Ohio, followed by twenty years as a full professor in the Department of Statistics at Iowa State University. He has extensive experience in linear statistical models, which is an area of statistics that makes heavy use of matrix algebra, and has taught graduate-level courses on that topic. He has authored over 70 research articles. His work has been recognized by his election as a Fellow of the American Statistical Association and the Institute of Mathematical Statistics and as a member of the International Statistical Institute. He has served as an associate editor of Biometrics and of the Journal of the American Statistical Association. ISBN 0-387-95318-3 www.springer-ny.com