/
ISBN: 1520-9121
Похожие
Текст
STUDENT MATHEMATICAL LIBRARY
Volume 44
A (Terse)
Introduction to
Linear Algebra
Yitzhak Katznelson
Yonatan R. Katznelson
^sa4&.
«AM,
American Mathematical Society
Providence, Rhode Island
Editorial Board
Gerald B. Folland Brad G. Osgood
Robin Forman (Chair) Michael Starbird
2000 Mathematics Subject Classification. Primary 15-01.
The cover art is created by Noah Katznelson
and used with permission.
For additional information and updates on this book, visit
www.ams.org/bookpages/stml-44
Library of Congress Cataloging-in-Publication Data
Katznelson, Yitzhak, 1934-
A (terse) introduction to linear algebra / Yitzhak Katznelson, Yonatan R.
Katznelson.
p. cm. — (Student mathematical library, ISSN 1520-9121 ; v. 44)
Includes index.
ISBN 978-0-8218-4419-9 (alk. paper)
1. Algebras, Linear. I. Katznelson, Yonatan R., 1961- II. Title. HI. Title:
Introduction to linear algebra.
QA184.2.K38 2008
512'.5—dc22 2007060571
Copying and reprinting. Individual readers of this publication, and nonprofit
libraries acting for them, are permitted to make fair use of the material, such as to
copy a chapter for use in teaching or research. Permission is granted to quote brief
passages from this publication in reviews, provided the customary acknowledgment of
the source is given.
Republication, systematic copying, or multiple reproduction of any material in this
publication is permitted only under license from the American Mathematical Society.
Requests for such permission should be addressed to the Acquisitions Department,
American Mathematical Society, 201 Charles Street, Providence, Rhode Island 02904-
2294, USA. Requests can also be made by e-mail to reprint-permissionOams .org.
© 2008 Yitzhak Katznelson and Yonatan R. Katznelson
Printed in the United States of America.
@ The paper used in this book is acid-free and falls within the guidelines
established to ensure permanence and durability.
Visit the AMS home page at http://www.ams.org/
10 9 8 7 6 5 4 3 2 1 13 12 11 10 09 08
Contents
Preface ix
1 Vector Spaces 1
1.1 Groups and fields 1
1.2 Vector spaces 4
1.3 Linear dependence, bases, and dimension 14
1.4 Systems of linear equations 22
*1.5 Normed finite-dimensional linear spaces 32
2 Linear Operators and Matrices 35
2.1 Linear operators 35
2.2 Operator multiplication 39
2.3 Matrix multiplication 41
2.4 Matrices and operators 46
2.5 Kernel, range, nullity, and rank 51
*2.6 Operator norms 56
3 Duality of Vector Spaces 57
3.1 Linear functionals 57
3.2 The adjoint 62
4 Determinants 65
4.1 Permutations 65
4.2 Multilinear maps 69
4.3 Alternating n-forms 74
4.4 Determinant of an operator 76
4.5 Determinant of a matrix 79
ν
vi Contents
5 Invariant Subspaces 85
5.1 The characteristic polynomial 85
5.2 Invariant subspaces 88
5.3 The minimal polynomial 93
6 Inner-Product Spaces 103
6.1 Inner products 103
6.2 Duality and the adjoint Ill
6.3 Self-adjoint operators 113
6.4 Normal operators 119
6.5 Unitary and orthogonal operators 121
*6.6 Positive definite operators 127
*6.7 Polar decomposition 128
*6.8 Contractions and unitary dilations 132
7 Structure Theorems 135
7.1 Reducing subspaces 135
7.2 Semisimple systems 142
7.3 Nilpotent operators 147
7.4 The Jordan canonical form 151
*7.5 The cyclic decomposition, general case 152
*7.6 The Jordan canonical form, general case 156
8 Additional Topics 159
8.1 Functions of an operator 159
8.2 Quadratic forms 162
8.3 Perron-Frobenius theory 166
8.4 Stochastic matrices 178
8.5 Representation of finite groups 180
A Appendix 187
A.l Equivalence relations-partitions 187
A.2 Maps 188
A.3 Groups 189
*A.4 Group actions 194
A.5 Rings and algebras 196
Contents
VII
A.6 Polynomials 201
Index 211
Symbols 215
This page intentionally left blank
Preface
This book is a presentation of the elements of Linear Algebra that
every mathematician, and everyone who uses mathematics, should
know. It covers the core material, from the basic notion of a finite-
dimensional vector space over a general field, to the canonical forms
of linear operators and their matrices, obtained by the decomposition
of a general linear system into the direct sum of cyclic systems. Along
the way it covers such key topics as: systems of linear equations,
linear operators and matrices, determinants, duality, inner products and
the spectral theory of operators on inner-product spaces. We conclude
with a selection of additional topics, indicating some of the directions
in which the core material can be applied and developed.
In its mathematical prerequisites the book is elementary, in the
sense that no previous knowledge of linear algebra is assumed. It is
self-contained, and includes an appendix that provides all the
necessary background material: the very basic properties of groups, rings,
and of the algebra of polynomials over a field. The book is intended,
however, for readers with some mathematical maturity and readiness
to deal with abstraction and formal reasoning. It is appropriate for an
advanced undergraduate course.
As the title implies, the style of the book is somewhat terse. We
mean this in two senses.
First, we focus with few digressions on the principal ideas and
results of linear algebra qua linear algebra. The book contains fewer
routine numerical examples than do many other texts, and offers almost
no interspersed applications to other fields; these should be adapted
to the readership and, if the book is used in a course, provided by the
teacher.
IX
χ
Preface
Second, the writing itself tends to be concise and to the point, to
the extent that some of the proofs might be better described as detailed
lists of hints. This is intentional—we believe that students learn more
by having to fill in some details themselves.
Besides its style, this book differs from many other texts on the
subject in that we try to present the main ideas, whenever possible, in
the context of vector spaces over a general field, F, rather than
assuming the underlying field to be R or C. Inner-product spaces, along with
the naturally associated classes of self-adjoint, normal, and unitary (or
orthogonal) operators, are introduced later than in many books, and
the spectral theorems for these operators, besides being fundamentally
important on their own, also serve here to pave the way for the notions
of reducing and semisimplicity and, eventually, to the general structure
theorems—the Jordan form, when the underlying field is algebraically
closed, and the corresponding form over general fields.
The text consists of eight chapters and an appendix. These are
divided into sections, and further into subsections. Definitions,
propositions, examples, etc., are numbered according to the subsection in
which they appear, and no subsection has more than one object
(definition, theorem, etc.) of each kind. For example, Lemma 1.3.5 is the
lemma appearing in subsection 1.3.5, and Theorem 1.3.5 is the
theorem appearing in the same subsection. References to the appendix
have the form A.x.y (for subsection у of section x, in the appendix).
Exercises appear at the end of sections, and are numbered
accordingly, e.g., exercise ex3.1.2 is the second exercise of section 3.1.
Starred sections, subsections, and exercises contain material that
can be skipped on first reading. Several of these sections, as well as
parts of the additional topics (in Chapter 8), require some familiarity
with basic analysis, e.g., concepts like convergence and continuity.
Chapter 1
Vector Spaces
1.1 Groups and fields
Vector spaces are defined over fields, and the definitions of both
fields and vector spaces depend on the notion of a commutative group.
Groups and fields are reviewed in more detail in the appendix but, for
convenience, we include here their definitions along with some basic
examples.
1.1.1 Groups.
Definition: A group is a pair (G, *), where G is a set and * is
a binary operation (x,y) ь-> х*у, defined for all pairs (x,y) £ Gx G,
taking values in G, and satisfying the following conditions:
G-l The operation is associative: For all x,y,z^ G,
(x * y) * ζ = χ * (у * ζ).
G-2 There exists a unique element e G G, called the identity element
or the unit of G, such that e * χ = χ * e = χ for all χ G G.
G-3 For every χ e G there exists a unique element *_1, called the
inverse of *, such that x~x *x = x*x~[ —e.
A group (G,*) is abelian, or commutative, if x^y = y^x for дЛ\
χ and y. The group operation in a commutative group is often written
and referred to as addition, in which case the identity element is written
as 0, and the inverse of χ as —x.
When the group operation is written as multiplication, the
operation symbol * is typically written as a dot (i.e., χ у rather than x*y)
1
2
1. Vector Spaces
and is often omitted altogether. We also simplify the notation by
referring to the group, when the binary operation is assumed known, as G,
rather than ((?,*).
Examples:
a. (Z, +), the integers with standard addition.
b. (R\ {0},·), the nonzero real numbers, with standard
multiplication.
c. Sn = (S„, ·), the symmetric group on[l,...,«];«a positive integer.
The elements of S„ are all the permutations σ of the set [1,..., n],
i.e., the set of the bijections (1-1 maps) of [1,..., n] onto itself.
The group operation is composition: for σ, τ Ε S„ we define τ·σ
by: (τ · a)(j) = τ(σ(])) for all j in [l,w].
The first two examples are commutative; the third is not if η > 2.
1.1.2 Fields.
Definition: A field (F,+,·) is a set F endowed with two binary
operations, addition, (а,Ь) н—> a + b, and multiplication, (a,b) y-^ a-b,
(usually written simply as ab), such that:
F-l (F, +) is a commutative group, whose identity element is denoted
byO.
F-2 (F\ {0},·) is a commutative group, whose identity element is
denoted by 1. It is the multiplicative group of F.
F-3 Addition and multiplication are related by the distributive law.
a(b + c) = ab + ac.
Examples:
a. Q, the field of rational numbers.
b. M, the field of real numbers.
1.1. Groups and fields
3
c. C, the field of complex numbers.
d. Z2 denotes the field consisting of the two elements 0, 1, with
addition and multiplication defined mod 2 (so that 1 + 1 =0).
More generally, if ρ is a prime, the set Ър of residue classes mod
p, with addition and multiplication mod /?, is a field. (See exercise
exl.1.4.)
The fields Q, M, and С are familiar, and are the most commonly
used. Less familiar, yet very useful, are the finite fields Ζρ, mentioned
above. Also important are extensions of a given field, see A.5.5.
EXERCISES FOR SECTION 1.1
exl.1.1 Verify that Sn is not commutative if η > 2.
exl.1.2 Show that if F is a field, then 0 · a = 0 for all a e F, and if ab = 0,
then a = 0 or b = 0.
exl.1.3 Verify that Z3 = {0,1,2} is a field if addition and multiplication are
defined mod 3, i.e., we add and multiply as usual, and if the result is > 3,
subtract 3; thus 1 + 1 = 2 but 2 χ 2 = 1.
Why is Z4, defined similarly—as the set {0,1,2,3} with addition and
multiplication defined mod 4—not a field?
exl.1.4 Let ρ > 1 be a positive integer. Recall that two integers m, η are
congruent mod p, written n = m (mod /?), if η — m is divisible by p. This
is an equivalence relation (see appendix A.l). For m £ Z, denote by m the
coset (equivalence class) of m, that is, the set of all integers η such that n = m
(mod p).
a. Every integer is congruent mod ρ to one of the numbers [0,1,...,/?— 1].
In other words, there is a 1-1 correspondence between Zp, the set of cosets
(mod /?), and the integers [0,1,...,/?— 1].
b. As in subsection 1.2.4 above, we define the quotient ring Ър = Z/pZ
(both notations are common) as the space whose elements are the cosets
(mod p) in Z, and define addition and multiplication by: m + h = (m + n)
and rh'h = nvn. Prove that the addition and multiplication so defined are
associative, commutative, and satisfy the distributive law.
c. Prove that Zp, endowed with these operations, is a field if and only if ρ
is prime.
4
1. Vector Spaces
Hint: You may use the following fact: if ρ is a prime, and both η and m are
not divisible by p, then nm is not divisible by p. Show that this implies that
if η φ 0 in Zp, then {nm :meZp} covers all of Zp.
1.2 Vector spaces
1.2.1 Definition. A vector space У over a field ¥ is an abelian group
(V,+), for which a binary product, (α,ν) н^ αν, of F χ ^ into У is
defined, satisfying the following axioms for all a,b e¥ and u,v еУ:
VS1. lv = v,
VS2. (ab)v = a(bv)9
VS 3. (<z + fc)v = <zv + fcv and a(v + u) =av + au.
Other familiar properties may be derived from these. For example, for
every vEf, Ov = (1 - l)v = v- v = 0.
The elements of Ψ are usually referred to as vectors', the elements
of the underlying field as scalars.
Observe that in the equality (ab)v = a{bv) the multiplication (ab)
is within F; the others are products of a vector by a scalar. In (a + b) ν =
αν + bv, the addition on the left is the addition in F, while that on the
right is the addition in Ψ.
Most of the notions and results we discuss are valid for vector
spaces over arbitrary fields. When the underlying field does not need
to be specified, we denote it by the generic F.
Results that apply to vector spaces over specific fields, or to vector
spaces over fields satisfying some additional conditions, will be stated
explicitly in terms of the appropriate field or the additional conditions.
If the underlying field is R or C, then the vector space is called a
real vector space or a complex vector space, respectively.
Vector spaces may also have additional structures: geometric, such
as inner-product, which we study in Chapter 6; or algebraic, such as
multiplication, as discussed in A.5.6.
Examples: The following sets are all vector spaces over the indicated
fields.
1.2. Vector spaces
5
a. Fw, the space of all F-valued n-tuples [ 611 .... . 6ί^2 with addition and
multiplication by scalars defined by
[av...,an] + [bv...,bn] = [al+bv...,an + bn],
c[av...,an] = [cav...,can].
We may write the η-tuples as rows, as we did here, or as columns.
If we want to specify that the vectors are written as columns or as
rows, we write F" or F", respectively.
b. M(n, m; F), the space of all F-valued nxm matrices; that is, arrays
\aU a\m\
A a2\ a2m
with entries from F. We sometimes write A = [a-] when the
dimensions of the matrix are assumed known, to save space.
The addition and multiplication by scalars are again done entry
by entry. As a vector space, Ж{п,т\¥) is virtually identical with
We write Jt{n\¥) instead of JZ(n,n\¥), and if the underlying
field is either known implicitly, or assumed explicitly, then we
often write simply Μ{η,πί) or *Ж(п), as the case may be.
c. ¥[x], the space1 of all polynomials Σαηχη with coefficients from
F. Addition and multiplication by scalars are defined formally
either as the standard addition and multiplication of functions (of the
variable χ ), or by adding (and multiplying by scalars) the
corresponding coefficients. The two ways define the same operations.
More generally, the set F[jcj ,... ,jcJ of all polynomials in к
variables over F is a vector space (in fact—an algebra).
l¥[x] is an algebra over F, i.e., a vector space with the additional operation of
multiplication. See A.5.6.
6
1. Vector Spaces
d. Let X be a finite set. C(X) denotes the set of all complex-valued
functions on X, with the standard addition of functions, and
multiplication of functions by scalars.
e. The set CR([0,1]) of all continuous real-valued functions / on
[0,1], and the set C([0,1]) of all continuous complex-valued
functions / on [0,1], with the standard operations of addition of
functions and of multiplication of functions by scalars. CR([0,1]) is
a real vector space. C([0,1]) is naturally a complex vector space,
but becomes a real vector space if we limit the allowable scalars to
real numbers only.
/. The set C°° ([— 1,1]) of all infinitely differentiable real-valued
functions /on [—1,1], with the standard operations on functions is a
real vector space.
g. The set 3?N of 27T-periodic trigonometric polynomials of degree
bounded by N\ that is, the functions of the form Σι„ι<^ ^η^ηχ- The
underlying field may be С or R, and the operations are the standard
addition of functions and multiplication of functions by scalars.
Similarly, the space ^N M of trigonometric polynomials in two
variables of the form Σι„ι<^ \т\<ме^ПХ+ту^ w^^ ^е same °Рег~
ations and the same underlying field(s).
h. The set of complex-valued functions / on R that satisfy the
differential equation
3f"(x)-sinxf(x) + 2f(x) = 0,
with the standard operations on functions. If we are interested in
real-valued functions only, the underlying field is naturally R. If
we allow complex-valued functions we may choose either С or R
as the underlying field.
1.2.2 Isomorphism. The expression "virtually identical" in the
comparison above of ^(n,m\¥) with ¥mn is not a proper mathematical
term. The proper term here is isomorphic.
1.2. Vector spaces
7
Recall (or see section A.2 in the appendix) that a map φ: Χ y-^Y
from a set X to a set Υ is bijective if for every у G Г there is precisely
one* G X such that у = φ (χ). Bijective maps are invertible—the inverse
map is defined by:
(1.2.1) <P~l(y)=x if У = <Р(х).
Let Ψ and W be vector spaces (over the same field).
Definition: A map φ: У н^ ψ is //near if for all scalars β, b and
vectors v,,v2 G Ψ
(1.2.2) φ(βν1 +£ν2) = αφ(ν1) + £φ(ν2).
A map φ: У н—> #^ is an isomorphism if it is bof/? bijective and linear.
An automorphism is an isomorphism of a vector space Ψ onto itself.
If φ is an isomorphism of Ψ onto #^, then φ-1 is an isomorphism
of W onto Ψ. This can be seen as follows: by (1.2.2),
(1 2 3) φ_1^ΐ(Ρνΐ +β2Φν2) =<P~l (Φ(β1ν1 +β2ν2)) = β1ν1 +β2ν2
= β1φ"1(φν1) + β2φ-1(φν2),
and, as φ is surjective, every vector in W is equal to φν for some
We say that Ψ andW are isomorphic if there is an isomorphism φ
of У onto #^, and the fact that the inverse φ-1 is also an isomorphism
guarantees that the relation of being isomorphic is symmetric.
The identity map φ(χ) =x shows that the relation is reflexive and,
since the composition of isomorphisms is an isomorphism, see
exercise exl.2.2, the relation is also transitive. In other words, the relations
of being isomorphic is an equivalence relation (see A. 1.2).
1.2.3 Subspaces. A (vector) subspace of a vector space Ψ is a
subset that is closed under the operations of addition and multiplication
by scalars inherited from Ψ. In other words, Ψ С Ψ is a subspace if
for all scalars a- and vectors w G W, j = 1,2, the vectors alwl +a2w2
are in W.
8
1. Vector Spaces
Examples:
a. Solution-set of a system of homogeneous linear equations.
Here Ψ = F". Given the scalars af.., 1 < i < k, 1 < j < n, we
consider the solution-set of the system of к homogeneous linear
equations
η
(1.2.4) Σβ/Λ = 0> i=h---,k.
This is the set of all η-tuples, [xx,... ,xn] G Fw (thought of as
vectors), for which all к equations are satisfied. If both [x{,... ,xn] and
[)Ί > · · · j^/i] are solutions of (1.2.4), and β and b are scalars, then for
each /,
η η η
£ aij{axj + ftyy.) = a £ ад + ft £ ад = 0.
It follows that the solution-set of (1.2.4) is a subspace of Fw.
ft. In F[jc], the space F^[x] of polynomials Y^anxn of degree < N.
While F[jc] is an algebra, ¥N[x] is not an algebra; why?
c. In the space C°° (R) of all infinitely differentiable, real-valued
functions / on R with the standard operations, the set of functions /
that satisfy the differential equation
f'(x)-5f(x) + 2f(x)-f(x) = 0.
If we consider complex-valued solutions, then the field of scalars
can be R or С
d. Subspaces of M(n)\
The set of diagonal matrices—the nxn matrices with zero entries
off the main diagonal, i.e., af. · = 0 for / φ j.
The set of symmetric matrices—the nxn matrices whose entries
satisfy α· · = α ··.
1.2. Vector spaces
9
The set of skew-symmetric matrices—the nxn matrices whose
entries satisfy atj = —α β.
The set of lower triangular matrices—the nxn matrices with zero
entries above the main diagonal, i.e., α· ■ = 0 for / < j.
Similarly, the set of upper triangular matrices, i.e., those for which
a- · = 0 for i > /.
ij J
2
Remark: In view of the isomorphism between Ж(п) and ¥n ,
these subspaces can be viewed as special cases of example a.
e. Intersection of subspaces: If W- are subspaces of a space У, j G 7,
(the index set J can be finite or infinite), then Π Ж is a subspace
ofr.
/. The sum of subspaces: If Ж, j G 7, are subspaces of У, their sum
is the set:
2X.= |J{v: v=£v;, VjeWj},
JxcJ J£JX
where the union extends over the collection of finite subsets of J.
Don't confuse the sum of subspaces with the union of subspaces,
which is seldom a subspace, see exercises exl.2.5 and exl.2.4.
g. The span of a subset: Given Ε С У, a linear combination of
elements of Ε is a f/n/te sum of the form Σβ,ν; ? β, £ i*\ v, £ £· The
set of all the linear combinations of elements of Ε is a subspace of
У, called the span of Ε and denoted by span[£].
1.2.4 Quotient spaces. A subspace W of a vector space У defines
an equivalence relation (see the appendix, section A.l) in У:
(1.2.5) x = y {modW) if x-yeW.
In order to establish that this is indeed an equivalence relation, we need
to verify that it is reflexive, symmetric, and transitive:
a. reflexive: x = x, since χ — χ = 0 G #^,
10
1. Vector Spaces
b. symmetric: χ = у <=> у = χ, since χ — у Ε W if and only if
y-x= —(x—y) Ε W,
c. transitive: If both χ = у and y = z, then χ = ζ. This follows from: if
both (x—y) and (y — z) are in W, thensoisx — z= (x—y) + (y — z).
This equivalence relation partitions Ψ into equivalence classes,
called the cosets of Ψ in ^. For χ e У, the coset of Ψ that
contains χ is the set χ = χ + #^ = {ν = χ + w : w Ε #^}—the "translate"
of #^ by x.
We define the quotient space У/W to be the space whose elements
are the cosets of W in Ψ, and the operations of addition and
multiplication by scalars are defined as follows. If χ = χ + W and у = y + W
are cosets, and α Ε F, then
(1.2.6) x + y = x + y + W = x + y and αχ = αχ.
The definition needs justification. We defined the sum of two
cosets by taking a representative element from each, and taking the
coset that contains their sum as the sum of the cosets. We need to show
that the result is well defined, i.e., that it does not depend on the choice
of the representatives in the cosets. In other words, we need to verify
that if χ = χχ (mod W) and у = yx (mod W), then x + y = xx+yx
(mod W). Now, if χ = xx + w, у = yx + w' with w, W Ε W, then
x + y = xx + w + yx + w' = xx +yx + w + v/, and since w + м/ Ε #^,
we have x + y = xx + yx (mod W).
The definition of ox is justified similarly: if χ = χχ (mod W), then
αχ —α*! = a(x — x{), and since #^ is a subspace, it is closed under
multiplication by scalars, α(χ—χχ) Ε W, and αχ ξ oxj (mod W) .
1.2.5 Direct sums. If Ψχ,..., Ук are vector spaces over F, f/?e (formal)
direct sum
к
ι
1.2. Vector spaces
11
is the set {(vj,..., v^) : ν · G ^} in which addition and multiplication
by scalars are defined by:
(vv...,vk) + (uv...,uk) = (vx+uv...,vk + uk),
α(ν1>···>ν*)= (avV">avk)·
The direct sum of vector spaces is clearly a vector space.
In the case that У^...,Ук are all subspaces of the same vector
space У, we defined their sum, Σ^ = {ν : ν = Y!j=\ v/> vj £ ^/}> *п
example/of 1.2.3.
The "natural" map of Ух θ · · · θ Ук into Ψλ Η \-Ук, defined by
(1.2.7) Φ((ν„...,ν,)) = Σ>7.,
ι
is clearly linear and surjective. It is an isomorphism when the sub-
spaces are independent.
Definition: The subspaces У-, j = 1,... ,&, of a vector space У
are independent if Σ ν = 0 with ν G У: implies that ν = 0 for all j.2
Proposition. Let У: be subspaces of У. The map Φ defined by (1.2.7)
is an isomorphism if and only if the subspaces are independent.
PROOF: Φ is clearly linear and surjective. It is injective if and only if
every vector in the range has a unique preimage, that is, if
(1.2.8) v), v] G У j and v'{ + ■ · · + v£ = v[ + · · ■ + v'k
implies that v" = v'· for every j. Subtracting and writing ν = v'j — Vy,
(1.2.8) is equivalent to: £v = 0 with ν G У у The subspaces are
independent if and only if this implies that ν = 0 for all j. <
In view of the proposition, we refer to the sum £ Ж of independent
subspaces of a vector space as their direct sum, and write 0 W- instead
of 150}.
If у = fy φ ψ, we refer to <% as a complement of Ψ in У, and
vice versa.
2Properly speaking: the set {УЛ is independent.
12
1. Vector Spaces
1.2.6 Tensor products. Let Ψ and fy be vector spaces over F. Let
Ψ ® % be the set of all the (finite) formal sums Σα.; ν · ® и ·, where
я · G F, ν · G ^ and и · G ^. We define addition formally by
jeJ{ jeJ2 jeJ{uJ2
and we define multiplication by scalars by
a ^ a j ν · ® w; = ^(ββ ·) ν · ® w;.
With these definitions, У ® ^ is a vector space over F.
The tensor product У ® ^ is, by definition, the quotient of У ® ^
by the subspace [Ψ ® ^/}0 spanned by the elements of the form
a. (vj +v2)®w — (vj ®w + v2®w),
(1.2.9) b. v®(m1 +w2) — (v^Mj + v®w2),
с a (v® и) — (αν)® и, (av)®u — v®(au),
for all ν, ν · G ^, и, и,- G ^, and α G F.
In other words, У ® <$/ is the space of formal sums Σβ/v/ ® w;
modulo the equivalence relation generated by:
a. (vl-\-v2)®u = vl®u-\-v2® w,
(1.2.10) b. v®(ul-\-u2) = v®ul-\-v®u2^
c. β (ν ® u) = (αν) ®u = v® (au).
Example: If Ψ = ¥[x] and fy = F[y], and we define the map Φ of
Г(§И^ intoF[x,y] by:
(1.2.11) Ф: £flyp/^)®iy(y)^£fly.py.(^(y)GF[^y],
then all the elements of [Ψ ® ^/}0 are mapped to zero, so that all the
elements in an equivalence class modulo [Ψ ® $/}0 are mapped to the
same polynomial. For example, every formal sum in Ψ ® ty/ that is
1.2. Vector spaces
13
equivalent to p(x) ®q{y) is mapped to p(x)q(y). It follows that Φ
induces a map of the quotient space Ψ ® % onto F[jc,y]
(1.2.12) Φ: £>/*) ® ^.(y) -> J>;-(*)<?;()0·
7 J
Φ is an isomorphism, and the spaces У ® tf/ and ¥[x,y] are
isomorphic.
EXERCISES FOR SECTION 1.2
exl.2.1 Verify that R is a vector space over Q, and that С is a vector space
over either Q or R.
exl.2.2 Let ^,y = 1,2,3, be vector spaces over the same field F. Let
<Pj : У[^У2 anc* Ф2 · ^2 и ^з ^е isomorphisms. Prove that φ2φ1 is an
isomorphism of Ψλ onto Ύν Conclude that isomorphism is an equivalence
relation for vector spaces (defined over the same field F).
exl.2.3 Verify that the intersection of subspaces is a subspace.
exl.2.4 Let fy and Ψ be subspaces of a vector space У, and neither of them
contains the other. Show that <%£ U W is not a subspace.
Hint: Take и e % \ Ψ, w e Ψ \ <% and consider и + w.
exl.2.5 Verify that the sum of subspaces is a subspace, and prove that
£^.=span[U^].
exl.2.6 Check that for every Ε С Ψ, span[£] is a subspace of Ψ, and is
contained in every subspace that contains E.
exl.2.7 If Ψχ is a subspace of Ψ and φ is an isomorphism of Ψ onto #^, then
φψχ is a subspace of W.
exl.2.8 Describe all the complements in R2 of the subspace X = {(jc, 0): χ £
Μ}.
exl.2.9 Prove that two subspaces Ύλ and Ύ2 in a vector space are independent
if'^ny2 = {0}.
exl.2.10 Prove that the subspaces W-; с Ψ, j = 1,... ,N are independent if
and only if ψ. Π £¥ . /^ = {0} for all ;.
14
1. Vector Spaces
exl.2.11 Let <ft and Ψ be subspaces of a vector space Ψ\ then Щ is a
sub space of ^ + Ж, and ^ Π W is a sub space of ^. Prove that the quotient
spaces
(1.2.13) (^ + Ж)/^ and Ψ/(<%Γ\Ψ)
are isomorphic.
H/nf; A coset of <% in (^ + Ж) has the form w+ ^, and Wj + ^ = w2 + ^
if and only if wl — w2 e %.
*exl.2.12 Assuming that F is infinite,3 and Ύ is a vector space over F, show
that the union of a finite number of subspaces of Ψ, none of which contains
all the others, is not a subspace.
H/nf; Let ^, j = 1,... ,fc be the subspaces in question. Show that there is
no loss in generality in assuming that their union spans Ψ. Now you need
to show that []У- is not all of У. Show that there is no loss of generality
in assuming that Ψλ is not contained in the union of the others. Now take
Vi £ V\ \ υ#ι ^/> and w Φ V\ \ show that av\ + w G U ^/> « ^ F, for no more
than lvalues of a.
1.3 Linear dependence, bases, and dimension
Let Ψ be a vector space. A //near combination of vectors v],..., vk
is a sum of the form ν = Σβ,ν, with scalar coefficients а .
A linear combination is nontrivial if at least one of the coefficients
is not zero; otherwise it is trivial.
1.3.1 Recall that if Л is a subset of a vector space У, then span [A],
the set of all linear combinations of elements in A, is a subspace of У,
(see example/of 1.2.3).
Definition: The set А с У is a spanning set if spa η [A] = У.
1.3.2 Definition. The set А с ^ is linearly independent if for every
sequence {vp...^} of distinct vectors in A, the only vanishing linear
combination of the ν .'s is trivial; that is, if Σα ν = 0, then α . = 0 for
all;4
3lf F is finite, then ¥n = UvGF? span[v] is a finite union of subspaces.
4Independence is a property of the set A; however, we often say, by abuse of
language, that the vectors (in A) are independent.
1.3. Linear dependence, bases, and dimension
15
If the set A is finite, we enumerate its elements as vx,..., vw and
write the elements in its span as Σα]ν]· В У definition, independence
of Л means that the representation of ν = 0 as a linear combination of
elements from A is unique. Notice, however, that this implies that the
representation of every vector in span [A] is unique. In fact, the equality
Σ[ cijVj = Σ\ bjVj implies Σ\ (aj ~ bj)vj = 0 so that a- = b · for all j.
A vector ν is linearly dependent on a set A if it can be represented
as a linear combination of vectors from A, that is, if ν £ span [A].
1.3.3 A minimal spanning set is a spanning set such that no proper
subset thereof is spanning.
A maximal independent set is an independent set such that no set
that contains it properly is independent.
Lemma.
a. A minimal spanning set is independent.
b. A maximal independent set is spanning.
PROOF: a. Let A be a minimal spanning set. If Σβ,ν, = 0, with
distinct ν · £ A, and for some k, ak φ 0, then vk = — a~x Σ^αΐνΐ· This
permits the substitution of vk in any linear combination by the
combination of the other ν ,'s, and shows that vk is redundant: the span of
{v : j φ к} is the same as the original span, contradicting the
minimality assumption.
b. If В is independent and и is not in span[#], then the union
{u} U В is independent: otherwise there would exist {vp..., vj С В
and coefficients d and с ·, not all zero, such that dw + £c ν = 0.
If d t^O, then u= —ά~ιΣ€ΐνΐ and w would be in span[v1?... ,vj С
spa η [Б], contradicting the assumption и £ span [Б].
If d = 0, we have Lc,v, — 0 with some nonvanishing coefficients,
contradicting the assumption that В is independent. It follows that if
В is maximal independent, then и £ span[B] for every и £ Ψ\ in other
words: В is spanning. <
Definition: A basis for Ψ is a set В с Ψ which is both spanning
and independent. A vector space is finite-dimensional if it has a finite
basis.
16
1. Vector Spaces
Thus, {vj,..., vn} is a basis for У if and only if every ν еУ has
a unique representation as a linear combination of {vp...,v„}. This
representation, ν = Σβ,ν,·> is the expansion of ν relative to the basis
{Vp...,V„}.
By the lemma, a minimal spanning set is a basis, and a maximal
independent set is a basis.
1.3.4 Proposition. If У is finite-dimensional, then:
a. Every finite spanning set can be trimmed to a basis.
b. Every independent set can be expanded to a basis.
PROOF: a. Let {vy-}^=1 be a spanning set for У. Call a vector vl
inessential if it is linearly dependent on {v,·}'·"* , and essential
otherwise. This classification clearly depends on the order in which the vectors
are enumerated. Observe that an inessential vl is linearly dependent on
the essential vectors preceding it.
Remove the inessential vectors. Since every ν ■ is either essential
or linearly dependent on the preceding essential vectors, the essential
vectors span У and are independent, hence they form a basis.
b. Let {uM={ be independent, and let {еЛп-х be a basis for У.
Write Wj = и for j = 1,..., k, and w^. . = e for j = 1,..., n. The
sequence {w ·} contains the basis {еЛ and is therefore spanning. Now
remove, as in part a, the inessential vectors to obtain a basis, and
observe that the first к vectors, namely {u-}k-v are all essential, and
hence form part of the basis. <
Remarks: The statement of a and its proof are valid for infinite
spanning sequences as well. See also 1.3.8.
The argument of b is refined somewhat in the following
subsection.
Examples:
a. The standard basis for F": we write e- for the vector whose / th
entry is equal to 1 and all the other entries are zero. Then {ex,..., en}
1.3. Linear dependence, bases, and dimension
17
is a basis for Fw, and the unique representation of ν
terms of this basis is ν = Σα-e -.
.an.
in
b. The standard basis for M{n,ni)\ let £· denote the η χ m matrix
whose if th entry is 1 and all the other entries are zero. Then {Ε(Λ
is a basis for Λέ{η,ηί), and the expansion of A = [α·] G Ж{п,т)
is Л = ΣαυΕυ.
c. The space F[jc] does not have a finite basis. The infinite sequence
{•*"}~=o *s both linearly independent and spanning, that is, a basis.
As we see in the following subsection, the existence of an infinite
basis, even of an infinite independent set, precludes a finite basis
and the space is infinite-dimensional.
Notice, however, that the subspace ¥N[x], consisting of all the
polynomials of degree at most N, is finite-dimensional since the
set of N+ 1 vectors, {xn}„=0, is a basis.
1.3.5 Steinitz' lemma and the definition of dimension.
Lemma (Steinitz). If span[vv... ,v„] = У and {uv...,um} is
linearly independent in У, then the vectors ν can be (re)ordered so that,
for every к = 1,..., m, the sequence {u{,..., uk, vk+v ..., vn} spans Ψ.
In particular, m<n.
PROOF: Write ux =£fl-v-; this is possible since span[vl5... ,v„] = >/.
Reorder the v,'s, if necessary, to guarantee that a{ φ 0.
Now Vj =a~l(u{ —Y!j=2ajvj)> so that span [mj^,.. . ,v„] contains
every ν · and hence is equal to Ψ.
Continue inductively: assume that, having reordered v-'s, if
necessary, we have {uv...,uk, vk+v ..., vn} spans Ψ.
If к = m we are done. If к < т write
к п
(1.3.1) uk+i = Lajuj+ Σ Vr
7=1 j=M
18
1. Vector Spaces
Since {ul,..., um] is linearly independent, at least one of the
coefficients bj is not zero. Reordering, if necessary, the remaining v-'s, we
may assume that bk+] φ 0. Now rewrite (1.3.1):
к п
(1.3.2) vik+1 = -fei^1(£fly.iiy.-iiik+1+ £ bjVj),
j=\ j=k+2
sothatv^ G spanfwj,.. -,uk+x,vk+2,.. .,v„], and, once again, the span
is У. Repeating the step m times proves the lemma. м
Theorem. If{v],..., vn} and {uv..., um} are both bases, then m = n.
PROOF: Since {vv...,vn} is spanning and {uv...,um} is
independent, we have m<n. Reversing the roles we have η <m. 4
Part of Steinitz' lemma repeats part b of proposition 1.3.4: in a finite-
dimensional vector space, every independent set can be expanded to a
basis by adding, if necessary, elements from any given spanning set.
The additional information here, that any spanning set has at least as
many elements as any independent set, is the basis for the current
theorem, and enables the definition of dimension.
Recall that a vector space У is finite-dimensional if it has a finite
basis.
Definition: The dimension of a finite-dimensional vector space У,
denoted άιχηΨ, is the number of elements in any basis for Ψ. The
definition is unambiguous, since all bases have the same cardinality.
As you are asked to check in exercise exl.3.6 below, a subspace
W of a finite-dimensional space Ψ is finite-dimensional and, unless
Ψ = Ψ, the dimension dim Ψ of Ψ is strictly lower than а\тУ.
The codimension of a subspace W in a finite-dimensional space Ψ
is, by definition, а\тУ — dimW.
1.3.6 Assume that Ψ and W are finite-dimensional vector spaces
over F, and dim^ = dim#" = d. Let {Vj}dj=l and {wj}dj=] be bases
for ψ and Ψ respectively.
1.3. Linear dependence, bases, and dimension 19
Every element ν еУ has a unique representation as a linear
combination of the basis elements, so the map φ: У y-^W defined by:
d d
(1.3.3) (Ρ(Σβ;ν;)=Σβ;νν;'
1 1
is unambiguous. Since {wj} is a basis of W, the map φ is a bijection
onto W and it clearly satisfies condition (1.2.2); in other words, it is
an isomorphism. Conversely, if Ψ is finite-dimensional and φ is an
isomorphism of Ψ onto a vector space Ψ, then the φ-image of a basis
of Ψ is a basis of Ψ. This proves the following theorem.
Theorem. Two finite-dimensional vector spaces over the same field
are isomorphic if and only if they have the same dimension.
Remark: The definition (1.3.3) of the map φ can be viewed as a two-
step process. The first step assigns to each basis element ν · its image
w ·; the second "completes the definition by linearity": if φ is to be
linear and <p(v ·) = w ·, then linearity forces, and in turn is guaranteed by,
(1.3.3). The fact that {v ·} is a basis guarantees that φ is well defined,
and this independently of whether {w ·} is a basis or not. Adding the
assumption that {w ·} is a basis for W guarantees that φ is bijective,
and hence an isomorphism.
1.3.7 The following observation is sometimes useful.
Proposition. Let $/ and W be subspaces of an η-dimensional space
Ψ, and assume that dim У/ + dim Ψ > п. Then У/ Π Ψ ф {0}.
PROOF: Let {Uj}lj=l be a basis for 9/ and {ν^}^=] be a basis for Ψ.
Since / + m > n, the set {u;}ί·=] U {w,}7=i is linearly dependent, i.e.,
there exists a nontrivial vanishing linear combination
Lcjuj+Ldjwj = °-
If all the coefficients с were zero, we would have a vanishing non-
trivial combination of the basis elements {\ν№=ν which is ruled out.
Similarly not all the d .'s vanish. We now have in ty/ Π Ж the nontrivial
vector ν = £cj;Uj = —Y^djW·. <
20
1. Vector Spaces
1.3.8 Infinite-dimensional vector spaces. Many examples of
vector spaces, in many areas of mathematics, are infinite-dimensional, i.e.,
they do not have finite bases. Assuming the axiom of choice, it can be
shown that every spanning set in an arbitrary vector space У can be
trimmed down to a minimal spanning set, i.e., one that does not contain
a proper subset that is spanning. Likewise, every independent set is
contained in a maximal independent set, i.e., one that is not properly
contained in a larger linearly independent set.
A subset 88 which is either a maximal independent set or a minimal
spanning set is a basis in the sense that every ν G Ψ is equal to a unique
finite linear combination of elements of 88. A basis, so defined, is
called a Hamel basis for an infinite-dimensional space.
In a typical study involving infinite-dimensional vector spaces this
is irrelevant! Infinite-dimensional spaces (of interest) usually come
with a topology, which allows one to introduce the notion of
convergence. In this context, convergent infinite sums are allowed, and bases
are defined accordingly. The field devoted to the study of topological
vector spaces is Functional Analysis.
EXERCISES FOR SECTION 1.3
exl.3.1 Show that the set {v : 1 < j < k} is linearly dependent if and only if
Vj = 0, or there exists / G [2, k] such that vl is a linear combination of vectors
in{v;.:l<;</-l}.
exl.3.2 Let Ψ be a vector space, and Ψ С Ψ a subspace5. Let v, и G Ψ \ Ψ,
and assume that и G span[#^, v]. Prove that ν G span[#^, u].
exl.3.3 What is the dimension of C5 considered as a vector space over R?
exl.3.4 Is R finite-dimensional over Q?
exl.3.5 Let fy, Ψ be subspaces of a vector space Ύ, with fy Π Ψ = {0}.
Assume that {u{,..., uk} С Щ and {wx,..., wz} С W are linearly
independent sets. Prove that {ux,..., uk} U {wx,..., wt} is linearly independent.
exl.3.6 Let Ψ be finite-dimensional. Prove that every subspace W С Ύ is
finite-dimensional, and that dim^ < dim Ψ with equality only if Ψ = Ψ.
5У\Ж denotes the difference set {ν : ν e Ψ and ν φ Ψ).
1.3. Linear dependence, bases, and dimension
21
exl.3.7 Show that if Ψ is finite-dimensional, then every subspace Ψ С Ψ is
a direct summand, i.e., there is a subspace W С Ύ such that Ύ = W Θ #".
Show by example that #" need not be unique (see exercise exl.2.8).
exl.3.8 Let Ύ be a finite-dimensional vector space, and srf a collection of
subspaces of Ύ. Prove that there are one or more minimal elements in <£/, that
is W £ <£/ such that no element in <£/ is a proper subspace of #^.
Similarly, show that <£/ has one or more maximal elements, i.e., elements
that are not contained in any other element of szi.
Remark: If Ύ- are subspaces of Ύ such that none is contained in another,
then every Ύ- is both minimal and maximal.
exl.3.9 Assume that Ύ is «-dimensional, and let ^,7 = 1,...,& be sub-
spaces of Ύ such that, for I < j <k, ^+i is a proper subspace of W-. Prove
that к < п.
exl.3.10 Let У and W be finite-dimensional subspaces of a vector space.
Prove that У + Ψ and Ψ Π ^ are finite-dimensional and that
(1.3.4) dim(rn^)+dim(r + ^)=dimr + dim^.
exl.3.11 Repeat exercise exl.2.11 in the context of finite-dimensional spaces
as follows:
a. Choose a basis {еЛ for ^ilf;
b. Complete it to a basis for fy by adding the vectors {uk}\
c. Complete it also to a basis for W by adding the vectors {w;};
d. Check that {e.} U {uk} U {wj is a basis for <% + #";
e. Identify bases for the two quotient spaces invoved.
exl.3.12 Show that if Ψ-, j = 1,... ,£, are finite-dimensional subspaces of
a vector space У, then YJW- is finite-dimensional and dim^^ < £dim#^,
with equality if and only if the subspaces Ж are independent.
exl.3.13 Let Ύ be an «-dimensional vector space, and let Ψχ С Ύ be a
subspace of dimension m.
a. Prove that the quotient space У/Ух is finite-dimensional.
b. Let {vx,..., vm} be a basis for Ψχ and let {vpj,..., wk} be a basis for
Ψ/Ψχ. For ;G [1, к], let w · be an element of the coset w ·.
Prove: {v1,...,vw}u{w1,...,wik} is a basis for У. Hence k-\-m = n.
22
1. Vector Spaces
exl.3.14 Let Ύ be a real vector space, and let v1?..., vp £ Ψ be linearly
independent. Let r; = [a; 1?... ,a; ], 1 < / < s be linearly independent
vectors in W. Prove that the vectors ux = Σ^αι /vy> / = 1,..., J, are linearly
independent in У.
exl.3.15 Let У and <ft be finite-dimensional spaces over F. Prove that the
tensor product Ψ ® ^ is finite-dimensional. Specifically, if {^,-}y=1 is a basis
for Г, and {/J£Li is a basis for <й\ then Ц-0/J, 1 < ; < n, \<k<m,
is a basis for Ψ <g> ^, so that dim У 0 ^ = dim Уdim <%.
*exl.3.16 Assume that Ύ is «-dimensional vector space over an infinite field
F, and let {#^} be a finite collection of distinct га-dimensional subspaces.
a. Prove that no W.is contained in the union of the others.
b. Prove that there is a subspace f ci which is a complement of every
Ψ-.
Hint: See exercise exl.2.12.
*exl.3.17 Assume that any three of the five M3-vectors ν · = (* ·,yj,Zj), j =
1,..., 5, are linearly independent. Prove that the vectors
^j = (^yj^XjypXjZpyjZj)
are linearly independent in R6.
Hint: Find nonzero (a,b,c) such that ax- + £v + cz = 0 for j = 1,2. Find
nonzero (d,e,f) such that dx- + ey. + /zy· = 0 for j = 3,4. Observe (and use)
the fact
(ax5 + by5 + cz5)(<£x5 + ey5 +fz5)^ 0.
1.4 Systems of linear equations
How do we find out if a set {v ·}, j = 1,... ,m of column vectors
in F" is linearly dependent? How do we find out if a vector и belongs
tospan[v,,...,vm]?
Given the vectors ν,
«/-
, j = 1,..., m, and и ■
, we
express the conditions Σ·*,ν, — 0 f°r the first question, and Σ·*/ν/ = и
for the second, in terms of the coordinates.
1.4. Systems of linear equations
23
The first equation leads to the system of homogeneous linear
equations:
(1.4.1)
or,
(1.4.2)
^1 l^l
a2\X\ +
an\xl +
m
Laijxi
= o,
~r a j w-*-ra
"^ a2mXm
~r Q>nmXm
i=l,..
= 0
= 0
= 0
.,n.
y=i
The second equation gives the nonhomogeneous system:
m
(1.4.3) Y,aijxj = cv i=l,...,«.
Definition: The solution-set of a system of linear equations, such
as (1.4.2) or (1.4.3), is the set of all m-tuples (jCj,...,jcm) G ¥m for
which all η equations hold.
To answer the question of the dependence of the ν ,'s, we need to
determine if the solution-set of the system (1.4.2) is trivial or not, i.e., if
there are solutions other than (0,..., 0). To see if и G spanfvj,..., vw],
we need to know if the solution-set of (1.4.3) is empty or not.
In both cases we would like to identify the solution-set as
completely and as explicitly as possible.
1.4.1 Conversely, beginning with a homogeneous system like (1.4.2)
we can rewrite it as
(1A4)
un\.
Η \-xn
= 0
and use general properties of vector spaces to draw general
conclusions. Our first result depends only on dimension.
Theorem. A system of η homogeneous linear equations inm> η
unknowns has a nontrivial solution.
24
1. Vector Spaces
PROOF: The m columns in (1.4.4) are elements of the rc-dimensional
space F". If m > n, then they are dependent, so (1.4.4) has a nontrivial
solution. Μ
Remark: As noted in example a of subsection 1.2.3, the set of
solutions of a homogeneous system of equations form a subspace of Fw.
The theorem above shows that if the number of variables is greater
than the number of equations, then the dimension of the subspace of
solutions is at least 1. See exl.4.1 for a refinement.
Similarly, a nonhomogeneous system like (1.4.3) can be rewritten
in the form
(1.4.5)
un\.
Η Yxn
It is then clear that the system given by (1.4.3) has a solution if and
only if
the column
is in the span of the columns
aU
J= l,...,m.
1.4.2 The classical approach to solving systems of linear equations
is by Gaussian elimination—an algorithm for replacing the given
system by an equivalent system that can be solved easily. We need some
terminology:
Definition: The systems
(21)
(1.4.6)
are equivalent if they have the same solution-set (in Fw).
1.4. Systems of linear equations
25
The matrices
A =
221
22ra
and AflM^ =
221
Ъ
a£ra C£-
^1 wbJ L"*l
are called the coefficient matrix, or simply the mafr/x, and the augmented
matrix of the system (21). The augmented matrix is obtained from the
matrix by appending the column of values (i.e., the right-hand sides of
the equations in the system) as a (new) last column.
The augmented matrix contains all the information of the system
(21). Any к χ (m + 1) matrix is the augmented matrix of a system of
linear equations in m unknowns.
1.4.3 Row equivalence of matrices.
Definition: The row space of a matrix A e Jtik.m) is the sub-
space of Fw spanned by the rows of A. The dimension of this space is
called the row rank of A.
The matrices
(1A7)
221
Ak\
22m
гкт-
and
'11
hi
b
Ira
^2ra
lm·
are row equivalent if their rows span the same subspace of Fw; equiva-
lently: if each row of either matrix is a linear combination of the rows
of the other. Row equivalent matrices clearly have the same row rank.
Proposition. Two systems of linear equations in m unknowns
(21)
(®)
are equivalent if their respective augmented matrices are row
equivalent.
ra
i=l,.
i=l,.
.,*,
.,/,
26
1. Vector Spaces
PROOF: Assume that the augmented matrices are row equivalent.
If (xx,... ,xm) is a solution for system (21) and
(biV ' · ',bimA) = Σα«Άΐ> ' * · >fl*m>C*)>
then
m
Σ buxj = Σ auk%xj= Σ aukck = di
y'=l к J к
and (x{,... ,xm) is a solution for system (55).
1.4.4 Reduced-row-echelon form. We come now to Gauss-Jordan
elimination. The equivalent system that is easier to solve is obtained by
reducing the augmented matrix of the system to one in reduced-row-
echelon form.
Definition: A matrix A e Ji{k,m) is in reduced-row-echelon form
if the following conditions are satisfied:
rref-l The first q rows of A are linearly independent in Fw, and the
remaining k — q rows are zero.
rref-2 There are integers \ <l{ <l2<··· <lq<m such that for j < q,
the first nonzero entry in the /th row is 1, occuring in the /;'th
column.
rref-3 The entry 1 in row j is the only nonzero entry in the / column.
One can rephrase the last three conditions as: The / 4h columns, called
pivot columns, are the first q elements of the standard basis of Fj:; every
other column is a linear combination of the pivot columns that precede
it.
Theorem. Every matrix is row-equivalent to a matrix in reduced-
row-echelon form. Furthermore, the reduced-row-echelon form of a
matrix is unique.
PROOF: We describe an algorithm that uses elementary row operations
to reduce an arbitrary matrix A to a row-equivalent matrix in reduced-
row-echelon form.
1.4. Systems of linear equations
27
The elementary row operations are:
a. Reordering (i.e., permuting) the rows;
b. Multiplying a row by a nonzero constant;
с Adding a multiple of one row to another.
These operations do not change the span of the rows so that the row
equivalence class of the matrix is maintained. (We shall return later,
in exercise ex2.3.13, to express these operations as matrix
multiplications.)
If A = 0, there is nothing to prove, and we assume that A^O.
Denote the row rank of Л by q. Let l{ be the index of the first column
that is not zero.
Reorder the rows so that αχι φ 0, and multiply the first row by
For every j > 1, subtract the first row multiplied by а . l from the
/th row.
Now all the columns before lx are zero and column lx has 1 in the
first row, and zero elsewhere.
If the row rank q is 1, all the entries below the first row are now
zero and we are done. Otherwise let /2 be the index of the first column
that has a nonzero entry in a row beyond the first. Notice that /2 >
/1. Keep the first row in its place, reorder the remaining rows so that
α21 φ 0, and multiply the second row6 by a~j .
For every j φ 2, subtract the second row multiplied by a ■l from
the /th row.
Repeat the sequence of steps a total of q times. The first q rows,
rj,..., rq, are (now) independent: a combination Lc,r, has entry с in
the / 4h place, and can be zero only if с = 0 for all j.
If there is a nonzero entry beyond the current g'th row, necessarily
beyond the Z^'th column, we could continue and get a row independent
of the first q, contradicting the definition of q. Thus, after q steps, all
the rows beyond the g'th are zero.
'We keep referring to the entries of the successively modified matrix as a- .
28
1. Vector Spaces
Uniqueness of the reduced-row-echelon-form of a matrix is left as
exercise exl.4.4. <
Observe that the scalars used in the process belong to the smallest
field that contains all the coefficients of A.
1.4.5 If A and Aaug are the matrix and the augmented matrix of a
system (21) and we apply the algorithm of the previous subsection to both,
we observe that since the augmented matrix has the additional column
on the right-hand side, the first q (the row rank of A) steps in the
algorithm for either A or Aaug are identical. Having done q repetitions, A
is in reduced-row-echelon form, while Aaug may or may not be. If the
row rank of Aaug is q, then the algorithm for Aaug ends as well;
otherwise we have / +1 = m+ 1, and the reduced-row-echelon form for the
augmented matrix is the same as that of A but with an added row and
an added pivot column, both having 0 for all but the last entries, and
1 for the last entry. In the latter case, the system corresponding to the
row-reduced augmented matrix has as its last equation 0=1 and the
system has no solutions.
On the other hand, if the row rank of the augmented matrix is the
same as that of A, the reduced-row-echelon form of the augmented
matrix is an augmentation of the reduced-row-echelon form of A. In
this case we assign arbitrary values to the so-called free variables, i.e.,
the variables χ-, ιφΐρ j = 1,... ,#. We then move the corresponding
terms to the right-hand side and, writing С for their sum, we obtain
(1.4.8) xij=cj> 7 = 1, ···,?·
Theorem. A necessary and sufficient condition for the system (21) to
have solutions is that the row ranks of the matrix and of the augmented
matrix of the system be equal.
The discussion preceding the statement of the theorem not only
proves the theorem but offers a concrete way to solve the system. The
unknowns are now split into two groups, q pivot variables and m — q
free ones. We have "m — q degrees of freedom": the m — q free
unknowns become free parameters that can be assigned arbitrary values,
and these values determine the pivot unknowns uniquely.
1.4. Systems of linear equations
29
Remark: Notice that the split into pivot and free unknowns depends
on the specific definition of reduced-row-echelon form; counting the
columns in a different order may result in a different split, though the
number q of pivot variables would be the same, equal to the row rank of
A. For example, for the "system" of one equation with two unknowns
χ + у = 1, either χ or у can be chosen freely, thereby determining the
other.
Corollary. A linear system of η equations in η unknowns with matrix
A has solutions for all augmented matrices if and only if the only
solution of the corresponding homogeneous system is the trivial solution.
PROOF: The condition on the homogeneous system amounts to "the
rows of A are independent", and no added columns can increase the
row rank. <
1.4.6 Definition: The column space of a matrix A e Ж{к,т) is
the subspace of F^ spanned by the columns of A. The dimension of
this space is called the column rank of A.
Theorem. The column rank of a matrix A is equal to its row rank.
PROOF: Linear relations between columns of A are solutions of the
homogeneous system given by A. If В is row-equivalent to A, the
columns of A and В have the same set of linear relations (see
Proposition 1.4.3). In particular, if the row rank of A is q and В is in reduced-
row-echelon form, then the q pivot columns in В are independent, and
every other column is a linear combination of these. <
We refer to the common value of the row and column ranks of A
simply as the rank of A and denote it by ρ (A).
1.4.7 Definition. A submatrix of a matrix A is a matrix В obtained
by deleting from A some rows and some columns.7
В is a principal submatrix of a square matrix A if the set of indices
of the deleted columns is the same as that of the deleted rows.
We use the word some to mean a no η negative number of, that is, some or none.
1. Vector Spaces
A simple corollary of Theorem 1.4.6 is the following proposition.
Proposition. If A is a matrix and ρ (A) = k, then there is а к х к
submatrix В of A such that ρ (В) = к.
See exl.4.12 for a refinement.
EXERCISES FOR SECTION 1.4
exl.4.1 Show that if m > n, then the dimension of the space of solutions of
a homogeneous system of η linear equations in m variables is at least m — n.
exl.4.2 Prove Proposition 1.4.7, and show an example of a 3 χ 3 matrix of
rank 2 without a principal submatrix of rank 2. (So that the second part of
the theorem is false without some assumption of symmetry.)
exl.4.3 Identify the matrix A e <Ж(п) of row rank η that is in reduced-row-
echelon form.
exl.4.4 Show that if А, Б £ Ж(к,т) are both in reduced-row-echelon form,
then either A = B, or A and В are not row equivalent. Conclude that the
reduced-row-echelon form of a matrix is unique.
Hint: Consider the solution sets of the two homogeneous systems with
coefficient matrices А зала Б, respectively. Show that these must be different if
АфВ.
exl.4.5 A system of linear equations with rational coefficients that has a
solution in C, has a solution in Q. Equivalently, vectors in Qn that are linearly
dependent over С are rationally dependent.
Hint: The last sentence of subsection 1.4.4.
exl.4.6 A system of linear equations with rational coefficients, has the same
number of degrees of freedom over Q as it does over C.8
exl.4.7 A subset szi of a vector space Ψ is called an affine subspace if srf is
the translate of a subspace W of Ύ, i.e., srf = {v0 + w :w £ W}. We call the
subspace Ψ the corresponding subspace to szi. (Thus a line in У is a translate
of a one-dimensional subspace.)
Prove that a set szi С Ύ is an affine subspace if and only if Σα -и- £ srf for
all choices of ux,..., uk £ «й^, and scalars а , j = 1,..., к such that £a ■■ = 1.
See 1.4.5.
1.4. Systems of linear equations
31
exl.4.8 Let szi С Ύ be an affine subspace and u0 G srf. Prove that the set
&i — u0 = {u — u0 : и e £/} is the corresponding subspace of srf in Ψ. Show
that the corresponding subspace szi — u0 does not depend on the choice of u0
msrf.
exl.4.9 Show that the solution-set of a system of k linear equations in m
unknowns is an affine subspace of Fm. What is its corresponding subspace?
exl.4.10 A column ν ,· =
4\j
\i.
of a matrix A =
Mm
лктЛ
is called a
pivot column if j is the index of a pivot column in the reduced-row-echelon
form of A. Show that ν is a pivot column of Л if and only if it is linearly
independent of the columns v·, i < j.
11 ··· Ъл
exl.4.11 Denote by В =
\m
the reduced-row-echelon form of
the matrix A in the previous problem. Let Ιλ < /2,... be the indices of the
pivot columns in В and i the index of another column. Prove that
(1.4.9)
where, as in the previous problem, Vj,..., vw are the columns of A.
exl.4.12 Show that if the matrix A in Proposition 1.4.7 is symmetric or skew-
symmetric, then the submatrix В may be taken to be a principal submatrix.
exl.4.13 What is the reduced-row-echelon form of the 7 χ 6 matrix A, if its
columns С.·, j = 1,..., 6, satisfy the following conditions:
a. Cx φ 0;
b.
с.
d.
C2 = 3Ci;
C3 is not a (scalar) multiple of Cx;
C4=q+2C2 + 3C3;
Cs = 6Co;
f. C6 is not in the span of C2 and C3.
32
1. Vector Spaces
*exl.4.14 Given polynomials Px = ΣΆα]χ^ P2 = LobjxJ> and s = lbsjxi
of degrees n, ra, and I < n + m respectively, suppose that we want to find
polynomials qx = L™_1 с -x7' and q2 = Lq_1 djx^ sucn mat
(1.4.10) Piqi+P2q2=S.
The polynomial equation (1.4.10) is equivalent to the system of m-\-n linear
equations,
(1.4.11) £ ajck+ Σ brdt=st, / = 0,...,/i + m-l,
where the unknowns are the coefficients cm_l,..., c0 of qx, and dn_x,...,d0
of q2. The coefficient matrix for this system is known as the Sylvester matrix
of (P„P2).
Write the matrix of the system for η = 3 and m = 2.
*exl.4.15 The associated homogeneous system to the system (1.4.11)
corresponds to the case S = 0. Show that it has a nontrivial solution if and only
if Ρχ and P2 have a nontrivial common factor. (You may assume the unique
factorization theorem; see A.6.3 in the appendix.) What is the rank of the
Sylvester matrix if the degree of gcd(Px ,P2) is r?
* 1.5 Normed finite-dimensional linear spaces
1.5.1 A norm on a real or complex vector space Ψ is a nonnegative
function vh ||v|| that satisfies the conditions
a. Positivity: ||0|| = 0 and if ν φ 0 then || v|| > 0.
b. Homogeneity: ||av|| = |<z|||v|| for scalars α and vectors v.
с The triangle inequality: ||v + m|| < ||v|| + ||и||.
These properties guarantee that δ(ν,ίΐ) = ||v — u\ is a metric on
the space, the metric defined or induced by the norm; and with a metric
one can use tools and notions from point-set topology such as limits,
continuity, convergence, infinite series, etc.
A normed vector space is a vector space endowed with a norm.
Since R and С are complete metric spaces, so is every normed finite-
dimensional real or complex vector space.
*1.5. Normed finite-dimensional linear spaces 33
1.5.2 If Ψ and Ψ are isomorphic real or complex finite-dimensional
spaces and 5 is an isomorphism of Ψ onto W, then a norm || · \Ψ on Ψ
can be "carried back" to Ψ by defining ||v||r = ||5v||^. This implies
that all possible norms on a real η-dimensional space are copies of
norms on R", and all norms on a complex η-dimensional space are
copies of norms on Cn.
A finite-dimensional Ψ can be endowed with many different norms;
yet, all these norms are equivalent in the following sense (see exl.5.1):
Definition: The norms ||-1| x and ||||2 are equivalent if there is a
positive constant С such that for all vGf,
— Ill II ^ II II ^ /~*\ I 11
IMIi < 1MI2 ь 4ΜΙι·
Metrics δρ <52, defined by equivalent norms 11-11 x and ||||2, are
equivalent: for v, и G Ψ,
C~l δ j (v, u) < δ2 (ν, и) < C<5 j (ν, и),
which means that they define the same topology—the familiar
topology of W or Cn.
EXERCISES FOR SECTION 1.5
exl.5.1 Let У be an «-dimensional real or complex vector space, and let
ν = {vj,..., vn] be a basis for Ψ. Write
||Vfl-v-|| =У|а.| and IIV^-v-ll =max|a-|.
Prove:
a. ||-||v j and ||·||ν.οο are norms on Ψ, and
(1-5.1) ΙΙ·ΙΙν,-<ΙΙ·ΙΙν,ι<»ΙΙ·ΙΙν,-
b. If || · || is any norm on Ψ then, for all v€f,
(1.5.2) ||v||vJmax||v;||>||v||.
c. Let ||·|| , j = 1,2, be norms on У, зала δ- the induced metrics. Let
{vn}™=0 be a sequence in Ύ. Prove that if 5j(v„,v0) -^0, then <52(v„,v0) —»0.
34
1. Vector Spaces
d. Show that if || · || is an arbitrary norm on Ψ, then the function /: F" ■
(where F is the field of scalars of У, either R or C), defined by
f(au...,an) = \\£,ajVj
is continuous on F". Conclude that / attains a strictly positive minimum on
the set Β = {(αλ,... ,α„) : I|fly-| = 1} С F".
Hint: Show that the triangle inequality implies that |||v|| — ||и||| < ||v — w||,
then use part b and the compactness of B.
e. Conclude that all norms on Ψ are equivalent to || · || v l.
exl.5.2 Let {vn}™=0 be bounded in Ψ. Prove that Y^znvn converges for
every ζ such that \z\ < 1.
Hint: Prove that the partial sums form a Cauchy sequence in the metric
defined by the norm.
exl.5.3 Let Ψ be «-dimensional real or complex normed vector space. The
unit ball in Ψ is the set
Bx = {v^f:\\v\\<\}.
Prove that Б j is
a. Convex: If v, и £ B], 0 < а < 1, then αν + (1 —a)ueBv
b. Bounded: For every vGf, there exists a (positive) constant Я such that
cv^Bfor \c\ > Я.
c. Circularly symmetric, centered at 0; If ν e В and \a\ = 1 then av e B.
Notice that convexity and circular symmetry together imply that if ν £ В
and \a\ < 1 then av £ B.
exl.5.4 Let У be «-dimensional real or complex vector space, and let Б be a
bounded circularly symmetric convex set centered at 0. Define
||w|| = 'mf{a>0:a~lueB}.
Prove that this defines a norm on У, and the unit ball for this norm is the
given set B.
exl.5.5 Describe a norm || ||0 on M? such that the standard unit vectors have
norml while ||(1,1,1)||0< γ^.
Chapter 2
Linear Operators and Matrices
2.1 Linear operators
2.1.1 Let У and Ψ be vector spaces over the same field F. Recall
definition 1.2.2:
Definition: A map T: У —> Ψ is linear if for all vectors ν ■ Ε У
and scalars α ·,
(2.1.1) Τ(αινι + α2ν2) = αχΤνχ + α2Τν2.
Induction on the the number к of summands proves that equation
(2.1.1) implies
к к
(2.1.2) T(L*J*j) = L*JT*J
7=1 j=\
for any number of summands.
Linear maps are also called linear operators, linear transformations,
homomorphisms, etc. The adjective "linear" is sometimes assumed
implicitly. The term that we use most is operator, and we always mean
by that a linear operator.
Examples:
a. The identity map / of У onto itself, defined by Ι ν = ν for ν Ε У.
For Я Ε F the multiplication operator ν н—> λ ν is a linear operator
on У. We denote it by λΐ or simply by A.
Z>. Let У be the space of all continuous, 27T-periodic functions on the
line. For every jc0 define Tx , the translation byx0:
τχ0: /(*) ^ A, (*) = /(* - *o) ·
35
36
2. Linear Operators and Matrices
c. The transpose,
e.
(2.1.3) A =
which maps jot
~аП a\m
a2\ a2m
βη\ anm_
^ A* =
'(n,m;F) onto «y#(m,n;F).
*11
*12
Ira
*л1
d. Differentiation on C°° ([0, 1 ]) (the complex vector space of infinitely
differentiable complex-valued functions on [0, 1]):
(2.1.4)
ax
Integration on C([0,1]) (the complex vector space of continuous
complex-valued functions on [0, 1]),
(2.1.5) Int:/*- / f(x)dx,
Jo
is a linear map from C([0,1]) into С
/. If Ψ — Ψ 0 ^, then every v e У has a unique representation
ν = w + и with w G Ж, и G ^; w is ffte component of ν in W, and
и the component in <$/. The map л^: vnw, mapping each vector
on its component in W, is the identity on W and maps <$/ to {0}.
It is called the projection of Ψ on W along ty/.
The operator πχ is linear since, if ν = w + и and vx=wx+ux, then
«v + fevj = (aw -\-bwx) + (ам + feiij). Now aw-\-bwx G Ж, and
au-\-bux G ^, so that Kx(av-\-bvx) = anxv-\-bnxvx.
Similarly, 7T2: ν н-> и is called the projection of У on ^ along W.
The maps %x and 7T2 are referred to as the projections corresponding
to the direct sum decomposition.
g. Restricting a linear operator. Let Ψ and Ψ be vector spaces over F
and let Τ:
be linear. If Ψχ С У is a subspace, we define
the restriction Τψ of Τ to Ψχ by setting, for ν G Ψχ, 7r ν = 7 v. Then
Гг is a linear operator from Ψχ into Ж.
2.1. Linear operators
37
2.1.2 Let У and Ψ be vector spaces over the same field F. Assume
that У is finite-dimensional, and let {vj,..., vn} be a basis. For any
choice of vectors w{,..., wn in W, the map ν ■ н-> w ·, j = 1,..., η
extends uniquely to a linear operator Τ from У loW, defined by:
(2.1.6) T--LaJvj"Lajwr
In other words:
Theorem. A linear operator from a finite-dimensional space Ψ to a
space W is completely determined by the values it assigns to the
elements of a given basis off, and these values can be arbitrary vectors
ιηψ.
PROOF: The map T, as defined by (2.1.6), is clearly linear, and, by
(2.1.2), a linear operator that agrees with Τ on {Vj,..., vn} must agree
with it on the entire space. Μ
Consider, for example, linear operators from F£ into F™. We choose
a basis ν = {vj,..., vn} of F^ and observe that, given this choice and
given a matrix A = (a-1) G <y#(m,n,F), we can define an operator TA
by setting TAvt to be the /'th column of the matrix:
^%
so that ТА(£см) =
(2.1.7)
ТаЧ
*l,i
LlCiflm,i.
and hence TA is completely determined by the explicit choice of the
basis v, the matrix A, and the (implicit) choice of the standard basis in
κ·
2.1.3 Given vector spaces Ψ and W, we denote the space of all
linear operators from Ψ into Ψ by jSf (Ψ, Ψ). Another common notation
is ΗΟΜ{Ψ,W). The two most important cases in what follows are:
Ψ = Ψ, and Ψ = F, the field of scalars.
• When Ψ = У, we write jSf (Г) instead of JSf (Г,Г).
• When Ж is the underlying field, we refer to the linear maps as linear
functional or linear forms on Ψ. Instead of «5f (^, F) we write ^*, and
refer to it as the dual space of У (see Chapter 3).
38
2. Linear Operators and Matrices
2.1.4 Coordinates. If Ψ is a finite-dimensional space, every basis
ν = {vj,..., vn} of У defines an isomorphism Cv of Ψ onto ¥n by
(2.1.8) Cv: v = £fly-vy.
= Lajer
The vector Cv ν is the coordinates vector of ν relative to the basis
v. Notice that this is a special case of (2.1.6) above: we map the basis
elements ν ■ on the corresponding elements e ■ of the standard basis,
and extend by linearity.
2.1.5 J£(Y,W) as a vector space. We define the sum of linear
maps T,S e Л?(У, W) and the multiple of a linear map by a scalar, as
follows: for every vE/,
(2.1.9) (T + S)v = Tv + Sv, (αΤ)ν = α(Τν).
Observe that {T + 5) and aT, as defined, are linear maps from Ψ to
W, that is, are elements of Jz?{Ψ,W). A straightforward verification,
left to the reader, shows that the addition and multiplication by scalars
as defined by (2.1.9) satisfy the requirements imposed on addition and
multiplication by scalars in the definition of a vector space. The
addition is associative, commutative, and distributive with respect to
multiplication by scalars. It follows that if У and W are vector spaces over
F, then, with addition and the multiplication by scalars as defined by
(2.1.9), &(У, Ψ) is a vector space over F.
Proposition. If both Ψ and W are finite-dimensional then so is
J5f(r,5^), andd\m^(y,W) = d\myd\mW.
The proof is left to the reader as exercise ex2.1.4 below.
EXERCISES FOR SECTION 2.1
ex2.1.1 Prove (verify) that the operator Τ defined in (2.1.7) is linear.
ex2.1.2 Show that if a set Л С Ψ is linearly dependent and Τ e а?(У, Ψ),
then ТА is linearly dependent in W.
2.2. Operator multiplication
39
ex2.1.3 Prove that an injective map Τ £ j£f (У, ^) is an isomorphism if and
only if it maps some basis of Ψ onto a basis of W, and this is the case if and
only if it maps every basis of Ύ onto a basis of W.
ex2.1.4 Let Ύ and #^ be finite-dimensional with bases {vl5...,vn} and
{wj,...,ww} respectively. Let <p·. G ^£{Ψ,Ψ) be defined by φ··ν· = w-
and φ· v^ = 0 for k φ i. Prove that {φ·;·: 1 < ι < и, 1 < j < m) is a basis for
ex2.1.5 Let Ψ and Ψ be finite-dimensional with bases {vl5...,v„} and
{wj,..., ww} respectively. Prove that «Sf (У, #^) is isomorphic to Л((т, η, F).
2.2 Operator multiplication
2.2.1 For Τ e JSf (Г, Ж) and 5 e jSf (5^, ^) we define the product
ST in &(У, <%) by composition, that is: (S7> = S(7V).
As defined, ST clearly maps Ύ into %; it is a //near operator since
(2.2.1) 5Γ(α1ν1 + α2ν2) ~ S^i^Vj + α2Γν2) ~ fljSTvj + a2STv2.
2.2.2 Of particular interest is the special case where Ψ — Ψ — %
so that T, 5, and 75 are all in &(У).
Proposition. With the product ST defined above, Jf(Y) is an
algebra1 over F.
PROOF: The claim is that the product is associative and distributive
(with the addition defined by (2.1.9)). That is: if R,SJ e JSf (У) then
R(ST) = (RS)T,
(2.2.2) V V
R(S + T)=RS + RT, and (R + S)T = RT+ ST.
This is a straightforward checking, left to the reader. <
The algebra Jz?(У) is not commutative unless dim Ψ — 1, in which case
it is simply the underlying field.
See the appendix, A.5.6.
40
2. Linear Operators and Matrices
The set of automorphisms, i.e., invertible elements in ££(Ψ\ is a
group under multiplication, the general linear group of У. It is denoted
GL(r).
Definition: The operators T,S G J£{Y) are conjugate if for some
R G GL(r),
(2.2.3) T = RSR~l.
The relation is symmetric (5 = R~lTR), reflexive, and transitive—it
is an equivalence relation.
2.2.3 Given an operator Τ G ^f (У), the powers T; of Γ are well
defined for all integers j > 1, and we define T° = I. Since we can take
linear combinations of the powers of T, we have P(T) well defined for
all polynomials Ρ G ¥[x\. Specifically, if P(x) = £^ then Ρ(Γ) =
We denote
(2.2.4) &>(T) = {P(T) : Ρ G F[*]}.
0*(T) is a subalgebra of Jz? {Ψ)\ it will be the main tool in
understanding the way in which Τ acts on У.
It is convenient to introduce the following terminology.
Definition: A linear system is a pair (V,T) where У is a vector
space and Τ e <5?(У). When we add adjectives, they apply in the
appropriate place, so that a finite-dimensional system is a system in which
Ψ is finite-dimensional, while an invertible system is one in which Τ is
invertible.
EXERCISES FOR SECTION 2.2
ex2.2.1 Give an example of operators T, S G ^(M2) such that Γ5 ^ ST.
ex2.2.2 Prove that, for any Τ G j£f (У), «^(Г) is a commutative subalgebra
of^f(r).
ex2.2.3 Let Τ G ^(R2) be defined by Tex = e2 and Te2 = ev Show that
&>{T) = {аТ + Ь:а,Ье Щ.2 What are the invertible elements in &>(T)1
2In this context, £ is a shorthand for the operator Ы.
2.3. Matrix multiplication
41
ex2.2.4 For Τ G jSf (Г) denote comm[r] = {S : S e JSf (У), 5Г = 7\S}, the
set of operators that commute with T. Prove that сотпл[Г] is a subalgebra of
ex2.2.5 Verify that GL(^) is in fact a group.
ex2.2.6 An element π G ^£(Ψ) is idempotent if ττ2 = π. Prove that an idem-
potent π is a projection onto its range, πΥ = {πν :v еУ}, along its kernel,
кег(тг) = {ν : πν = 0}.
2.3 Matrix multiplication
2.3.1 The product r · с of a row r (the 1 χ η matrix r = [av...,an])
V
) is, by definition, the scalar
and a column с (the η χ 1 matrix с =
given by
(2.3.1) vc^ajbj.
Given A G M(l,m) and В G <y#(m,n), we define the product AB as
the / χ η matrix С whose entries c- are given by
(2.3.2) cy = rI.(A)-cy.(e) = £flflkfe
У
(r-(A) denotes the /'th row in A, and с .(В) denotes the /th column in
B).
Notice that the product is defined only when the number of columns
in A (the length of its rows) is the same as the number of rows in B,
(the height of its columns).
The product is associative: given A G M(l,rri), В G M{m,n), and
С G Ji{n,p), then AB G Ж {I,n) and the product (AB)C G Ji{l,p)
is well defined. Similarly, A(BC) is well defined and one checks that
A(BC) = (AB)C by verifying that the r, s entry in either is £г аг-Ь -fis.
The product is distributive: for A G <y#(/,m), β. G M{m,n),
(2.3.3) (Aj +Α2)(β1 + 5г) = A151 +A1#2 + A2#1 + A2#2>
and commutes with multiplication by scalars: (аА)В —А{аВ) — а(АВ).
42
2. Linear Operators and Matrices
Going back to the definition and observing that A i—► r-(A) and
Β η—> с -(B) are linear maps, we have
Proposition. The map (A,B) н->А#, ofJ%(l,m) x^(m,n) to M(l,n),
is linear in В for every fixed A, and linear in A for every fixed B.
2.3.2 Write the m χ η matrix A = {aij)\<i<m as a "single row of
columns",
\<j<n
A2\
Am\
hn
hn
mn_
=
~an~
a2\
-am\-
~a\i
a22
Pml-
~aln
a2n
&mn_
(с1(А),с2(А),...,си(А));
we have, for every η-column y,
л2\
Lwml
In"
2n
mn_
~У\~
У2
Уп.
= (с,(А),с2(А),...,си(А))
У\
У2
\Уп.
= Еул(а),
7=1
so that for у G FJi, every column Ay is a linear combination (with weights y.)
of the columns of A.
Similarly, we can write the η χ ρ matrix Б as a "single column of rows",
B-
21
[*21
where r,-(fi) is the row (biX,..., bi ) € ¥ξ.
If [*,,...,*„] €FJ, then
ynp\
r2(B)
Уп(В).
(2.3.4) [*„...,*„]
y21
'n\
4>
'rcpj
t| , . . . ,H,nj
r,(B)·
r2(£)
.r»(B)J
Σ*Λ(*).
t=l
Proposition. Assume A G Ж{т^п)у В G Μ {η,ρ); then every column
of the matrix AB is a linear combination of the columns of A, and every
row of AB is a linear combination of the rows ofB.
We leave the final verification as an exercise (ex2.3.2) to the reader.
2.3. Matrix multiplication
43
2.3.3 \il — m — n, matrix multiplication is a product within Ж{п).
Proposition. With the multiplication defined above, Ж{п) is an
algebra over F. The matrix I = In = (δ- k) = YJ[EU is the identity* element
in Ж{п).
The invertible elements in Ж{п) form a group under
multiplication, the general linear group GL(n,F).
Theorem. A matrix A £ Ж(п) is invertible if and only if ρ (A) = n.
PROOF: Proposition 2.3.2 guarantees that the row rank of BA is no
bigger than the row rank of A. If В A = /, the row rank of A is at least
the row rank of /, which is clearly n.
On the other hand, if A is row equivalent to /, then its reduced-
row-echelon form is /, and by exercise ex2.3.13 below, reduction to
reduced-row-echelon form amounts to multiplication on the left by a
matrix B, so that BA = /, i.e., A has a left inverse. This implies that A
is invertible (see Exercise ex2.3.15). Μ
Definition: The matrices А, В £ Ж (η) are conjugate (by GL(n,F))
if there exists С £ GL(n,F) such that
(2.3.5) A = CBC~l.
As for operators, this is an equivalence relation.
EXERCISES FOR SECTION 2.3
ex2.3.1 Let r be the 1 χ η matrix all of whose entries are 1, and с the η χ Ι
matrix all of whose entries are 1. Compute re and cr.
ex2.3.2 Prove Proposition 2.3.2.
ex2.3.3 A square matrix (α· ·) is diagonal if the entries off the diagonal are
all zero, i.e., i φ j => at· = 0.
3δ·£ is the Kronecker delta, equal to 1 if j = к, and to 0 otherwise; £· is the
matrix whose entries ak l are zero except that ai ■ = 1.
44
2. Linear Operators and Matrices
Prove: If Λ is a diagonal matrix with distinct entries on the diagonal, and
if Б is a matrix such that AB = BA, then В is diagonal.
ex2.3.4 Denote by E(n',iJ), 1 < i Φ j < n, the η χ η matrix obtained from
the identity by interchanging rows i and j, i.e.,
Let Л £ JZ(n,m) and Б £ <M{m,ri). Describe Ξ (n\iJ)A and BE (n;i,j).
ex2.3.5 Let σ be a permutation of [1,... ,и]. Let Ασ be the η χ и matrix
whose entries а · · are defined by
fl if/ = a(y),
(2.3.6) a.={ KJh
I 0 otherwise.
Show that Ασ has precisely one 1 in each row and in each column, and all of
its other entries are 0. Conversely, show that if Λ is a matrix with precisely
one 1 in each row and in each column, and all of its other entries are 0,
then A = Ασ for some permutation σ of [1,..., л]. Such matrices are called
permutation matrices.
What is the transpose (Ασ) of the permutation matrix Λσ?
ex2.3.6 Show that the map σ \-^ Ασ defined above is multiplicative, that
is: Αστ = ΑσΑτ (στ is defined by composition: от(у') = a(r(j)) for all
ex2.3.7 Show that every permutation matrix Ασ £ <M{n) is a product of
matrices E(n\iJ).
Hint: See 4.1.2.
ex2.3.8 Let Ασ G Ж[п) be a permutation matrix. Let В £ <Ж(п,т) and
С e <Л((т,п). Describe ΑσΒ and CA0.
ex2.3.9 Let В e <M{n) be skew-symmetric, σ G Sn. Prove that ΑσΒΑ~ι is
skew-symmetric.
Similarly, if Б is symmetric, then so is ΑσΒΑ^1.
ex2.3.10 Denote by Et-, 1 < ij < n, the η χ η matrix whose entries are all
zero except for the ij entry which is 1. Let A £ JZ(n,m) and В £ Ж{т,п).
Describe E^A and BEt·.
ex2.3.11 Describe an η χ η matrix K(i,c) such that if Б £ JZ(n,m), then
K(i, c)B is the matrix obtained from Б by multiplying its fth row by c.
2.3. Matrix multiplication
45
ex2.3.12 Let В £ ЛК(п,т). Describe a square matrix A(c, ij) such that
multiplying В by it on the appropriate side has the effect of replacing the fth
row in В by the sum of the f th row and с times the / th row. Do the same for
columns.
ex2.3.13 Show that each of the steps in the reduction of a matrix A to its row-
echelon form (see 1.4.4) can be accomplished by left-multiplication of A by
an appropriate matrix, so that the entire reduction to row-echelon form can
be accomplished by left-multiplication by an appropriate matrix. Conclude
that if the row rank of A £ Ж{п) is n, then A is left-invertible.
ex2.3.14 Let A e <M{n) be non-singular and let В = (A,/), the matrix
obtained by "augmenting" A by the identity matrix, that is, by adding to A the
columns of / in their given order as columns η + 1,..., In. Show that the
matrix obtained by reducing В to row-echelon form is (/,A-1).
ex2.3.15 Prove that if A e JK(n,m) and В G JK(m,l), then4 (ABf = B?A*.
Show that if A £ ^(n) has a left inverse, then A has a right inverse, and if
A has a right inverse, then A has a left inverse. Use the fact that A and A
have the same rank to show that if A has a left inverse, B, it also has a right
inverse, C, and since В = В (AC) = (BA)C = C, we have BA=AB = I and A
has an inverse.
Where does the fact that we deal with finite-dimensional spaces enter the
proof?
ex2.3.16 Prove that the operator P(x) ι—» xP(x) on ¥[x] is left-invertible, but
not right-invertible.
ex2.3.17 What are the ranks and the inverses (when they exist) of the
matrices
0 2 10]
117 1
2 2 2 2
0 5 0 0_
·>
11111"
0 2 2 11
2 12 12
0 5 0 9 1
0 5 0 0 7
1111
0 111
0 0 11
0 0 0 1
0 0 0 0
ex2.3.18 Denote An-
1 η
0 1
. Prove that AmAn = Am+n for all integers m, n.
For the notation, see 2.1, example с
46 2. Linear Operators and Matrices
ex2.3.19 Let A,£,C,D G ^(/i), and let S =
whose top left
Prove that <f2
G <JK(2ri) be the matrix
r B\
[c d\
whose top left quarter is a copy of A, the top right quarter a copy of B, etc.
~A2+BC AB + BD]
CA+DC CB + D2
2.4 Matrices and operators
2.4.1 We have seen (2.1.7) that every matrix A G M(m, n, F) defines
an operator 7^ G J2?(F",F™). A quick check shows that TA is the
operator of left-multiplication by A of the columns of F". Conversely,
given Τ G JSf (F£, F™), if we take A = A7 to be the m χ η matrix whose
columns are Те-, where {ev...,en} is the standard basis in Fw, we
have TA = T. In other words, the maps АиГд and Гн^Аг are each
the inverse of the other. Since both maps are clearly linear, we obtain
the following theorem.
Theorem. The map Τ н—> AT is an isomorphism ofJf(F",F™) onto
m,n).
2.4.2 Let Τ G ^f(F",Fw), 5 G JSf(Fm,Fz), and let AT G Л{т%п\
resp. As G «y#(/,m) be the corresponding matrices; then
ST G JSf (F*,^), ASAT G M{l,n\
and, since the multiplication in both sides is composition,
In particular, if η = m = /, we obtain
Theorem. The map Τ н-> Аг w шг algebra isomorphism of Jf (Fn)
onto Ж{п).
2.4.3 The special thing about F" and F™ is that they have "standard
bases". The correspondence Τ <-> AT (or A «-» 7^) uses these bases
implicitly.
Consider now general finite-dimensional vector spaces У and W.
Let Τ G &{Ψ,Ψ\ and let v = {v1,...,v„}bea basis for Ψ and w a
basis for Ж.
2.4. Matrices and operators
47
Define the scalars tk- by the expansion: TV · = ΣΤ=ι tk]wki ^еп' ^or
any vector ν = £c ,v ·, we have
(2.4.2) 7v = J^CjTvj = ZLcAjwk = l(lcA,,H
J к к j
Given the bases {vp...,v„} and {w1,...,ww}, the full information
about Τ is contained in the matrix
= [CwTvv...,CwTv„].
The "coordinates operators", Cw, assign to each vector in W the column
of its coordinates with respect to the basis w; see (2.1.8).
Given the bases ν and w, and the matrix AT v w, the operator Τ is
explicitly defined by (2.4.2), which says: The column of w-coordinates
of Tv is obtained by multiplying the column of ν-coordinates of ν on the
left by the matrix AT v w; that is:
(2.4.4) Cw7v = A7vwCvv.
Let A G Ж{т,п), and denote by Sv the vector in W whose
coordinates with respect to w are given by the column A Cv v. So defined,
5 is clearly a linear operator in «if (Ψ, W) and ASyyf= A. This gives:
Theorem. Given vector spaces Ψ and W with bases v = {v1,...,v„}
and w = {w,,..., wm} repectively, the map Τ y-^ ATyyv is an
isomorphism of ^{У ,W) onto Ji{m,n).
When Ψ = Ψ and w = v, we write ATy instead of ATyy. The
reader should check that, just as in subsection 2.4.2, the isomorphism
Τ н^ AT y is an algebra isomorphism of «if (У) onto Ji{n).
2.4.4 Change of basis. Assume now that Ψ = Ψ, and that ν and w
are arbitrary bases. The v-coordinates of a vector ν are given by Cv v,
and the w-coordinates of ν by Cw v. If we are given the v-coordinates
(2.4.3)
7\v,w
Ml
Χ2\
m\
l\n
hn
48
2. Linear Operators and Matrices
of a vector v, say χ = Cv v, and we need the w-coordinates of v, we
observe that ν = С"1 χ, and hence Cw ν = Cw C"1 x. In other words, the
operator
(2.4.5) cW;V = cwc;1
on ¥n assigns to the (column of) v-coordinates of a vector vGf
the column of its w-coordinates. The opeator C"1 identifies the
vector from its v-coordinates, and Cw assigns to the identified vector its
w-coordinates; the space У remains in the background. Notice that
Suppose that we have the matrix A7w of an operator Τ G &(У)
relative to a basis w, and we want to have the matrix AT v of the same
operator 7\ but relative to a basis v. Claim:
(2.4.6) AT v = Cv.w^7)WCw.v;
CW;V assigns to the v-coordinates of a vector ν G У its w-coordinates;
AT w replaces the w-coordinates of ν by those of Tv\ Cv,w identifies Τ ν
from its w-coordinates, and produces its v-coordinates.
2.4.5 How special are the matrices (operators) CW)V? They are clearly
invertible, and that is a complete characterization.
Proposition. Given a basis v/ = {w],...,wn}ofy,the жар ν ι > Cw у
is a bijection of the set of bases у of У onto GL(n,F).
PROOF: Since Cw is non-singular, the equality CW)Vl = Cw.v2 implies
C"1 = C"1, and since C"1 maps the elements of the standard basis
of ¥n onto the corresponding elements in vp and C"1 maps the same
vectors onto the corresponding elements in v2, we have \x = v2. This
proves the injectivity.
To prove the surjectivity, let 5 G GL(n,F) be arbitrary. We shall
exhibit a base ν such that 5 = Cwv. By definition, Cww = e · (recall
that {ex,...,en} is the standard basis for Fw). Define the vectors ν · by
the condition: Cw ν ■ = Se;, that is, ν ■ is the vector whose w-coordinates
are given by the /th column of the matrix 5. As 5 is non-singular the
ν 's are linearly independent, hence form a basis ν of У.
2.4. Matrices and operators
49
For all j we have ν ■ = Cv l e ■ and Cw.ve, = Cw ν ■ = Se . This
proves that 5 = CW)V. ^
2.4.6 Similarity (matrices). By definition, the matrices Bx and B2
are similar if they represent the same operator 7 in terms of (possibly)
different bases, that is, B] = AT v and Б2 = A7 w.
If #j and B2 are similar, they are related by (2.4.6). By Proposition
2.4.5 we have
Proposition. The matrices Bx and B2 are similar if and only if there
exists С G GL(n,F) such that
(2.4.7) B]=C~]B2C, or equivalent^, CB]=B2C.
In other words, similarity for A and Б is synonymous to
conjugation (under GL(n,F)). We shall see later (see exercises ex4.5.18 and
ex7.5.4) that if there exists such С for (2.4.7) with entries in some field
extension of F, then there exists one in <y#(n,F).
A matrix is diagonalizable if it is similar to a diagonal matrix.
2.4.7 Similarity (operators). The operators 5, Τ G Л£{У) are said
to be similar if they are conjugate under GL(^), that is, if there is an
operator ReGL(Y) such that
(2.4.8) T = RSR~l or, equivalently, RS = TR.
An operator is diagonalizable if its matrix (relative to some basis) is
diagonalizable5. Observe that the matrix ATy of Τ relative to a basis
ν = {vj,..., vn} is diagonal if and only if Τν· = λ ·ν ·, where A, is the
/th entry on the diagonal of AT v.
2.4.8 Similarity (linear systems). A linear system is a pair (У,Т)
where У is a vector space and Τ € Л?(У).
5Because of Theorem 2.4.2, we apply to matrices the language and terminology
introduced for operators, and vice versa.
50 2. Linear Operators and Matrices
Definition: The systems (yvTx) and (^>,Γ2) are similar (other
terms used: conjugate, and isomorphic) if there is an isomorphism Ψ
of Ψχ onto Ψ2 such that
(2.4.9) ψη = 72Ψ.
The condition (2.4.9) is often described by stating that the following is
a commutative diagram:
Ψχ —^ Ψχ
(2.4.10)
Example: Right-multiplication by matrices in <y#(n) defines linear
operators on F". That is, A G */#(w) defines the linear map RA by:
(2.4.11) RA[av...,an) = [a{,...,an]A.
If we take the transpose on both sides, we obtain
(2.4.12) [RA[av...,an}}Tr=ATr[a{,...,anf\
Let Ψ denote the isomorphism of F" onto F" given by transposition.
The equality (2.4.12) is identical with (2.4.9) where Tx = RA, and T2 is
left-multiplication by ΑΉ on F". Thus, right-multiplication by A on FJ!
is similar to left-multiplication by A* on F".
EXERCISES FOR SECTION 2.4
ex2.4.1 Prove that S, Τ G ££ (У) are similar if and only if their matrices
(relative to any basis) are similar. An equivalent condition is: for any basis w
there is a basis ν such that AT v = As w.
ex2.4.2 Prove that two linear systems, (Ψ, Τ) and (W, S), are similar if and
only if dim Ψ = aimW and, if ν is a basis for Ψ and w a basis for W, then
AT v and A5 w, the matrices of Τ and 5 relative to these bases, are similar.
ex2.4.3 Let RA be defined on FJ by (2.4.11).
Ψ
Ψ
2.5. Kernel, range, nullity, and rank
51
a. What are the images of the elements of the standard basis in FJi?
b. Prove that the correspondence A <-> RA is an isomorphism of vector spaces.
Is it an algebra isomorphism?
ex2.4.4 Let ¥n[x] be the space of polynomials Σ^α-χΚ Let D be the
differentiation operator and Τ = 2D + /.
a. What is the matrix corresponding to Τ relative to the basis {х->}п-01
b. Verify that, if u- = LL;^' men {ujYj=q *s a basis, and find the matrix
corresponding to Τ relative to this basis.
ex2.4.5 Let 3?N be the complex vector space of 27T-periodic trigonometric
polynomials of degree < N\ that is, the space of functions / of the form
f(t) = Σ\η\<Ναη£ιηΐ- Let D : / \-^ f be the differentiation operator, and
T = D2 + ID. What is the matrix of Τ relative to the basis {eira}^<Nl
ex2.4.6 Let Ψ be «-dimensional, ν e У, and Τ G .Sf (У) such that {Tjv}nrl0
are linearly independent.
a. Verify: {Th}^ is a basis for У, and hence Γ"ν = Σ}Ζοα]τ)ν (for
appropriate coefficients a ·).
b. What is the matrix of Γ relative to the basis {Τ^ν}ηζ^Ί
ex2A.7 Let У be «-dimensional, v, w G У, and Г G .Sf(y).
a. Assume that the subspace Ψχ spanned by {Th}nr^ is ^-dimensional.
Prove that {Th}kr}Q is linearly independent (and hence is a basis for Ψχ).
b. With ν as above, assume that Ψ2, the span of {Tjv}nrl0 U {^;w}^, is
га-dimensional. Prove that {T^v}*"^ U {Tjw}m~k~l is linearly independent
(and hence is a basis for У2)·
ex2.4.8 Let A £ JZ(l,m). Prove that the map Г: Б »—» АВ of jtft{m,n) into
Ж(1,п) is a linear operator. In particular, if η = 1,<у#(т, 1) = F™, ./#(/, 1) =
F£ and Г e £?(¥™,Wlc). What is the relation between A and the matrix AT
defined in 2.4.3 (for the standard bases, and with η there replaced here by /)?
2.5 Kernel, range, nullity, and rank
2.5.1 Definition: The kernel of an operator Τ e Л?{У, Ψ) is the
set
ker(7) = {v e Г :Tv = 0}.
52
2. Linear Operators and Matrices
The range of Τ is the set
range(7) = ТУ = {w G Ψ : w = Tv for some ν G У}.
The kernel is also called the nullspace of T. An operator Τ is called
nonsingular if ker(T) = {0} and singular otherwise.
Proposition. Assume Τ e ^f(Y,W). Then ker(T) is α sub space of
У, and range(r) is a sub space ofW.
PROOF: If vp v2 G ker(T), then TX^Vj + α2ν2) =α{Γνλ +α2Τν2 = 0.
To show that range(r) is a subspace of Ж observe that if w = 7V-,
7 = 1,2, then tfjWj + #2w2 = ^(βινι +β2ν2)· ^
If Ψ is finite-dimensional and Τ G ££{?,W) then both ker(7)
and range(T) are finite-dimensional; the first since it is a subspace of
a finite-dimensional space, and the second as the image of one (since,
if {vp... ,v„} is a basis for У, {TVp... ,Tvn} spans range(T)).
We define ρ(Γ), the ran/c of 7\ as the dimension of range(r). We
define v(7), the nullity of 7\ as the dimension of ker(T).
Theorem (Rank and nullity). Let Ψ be finite-dimensional, W
arbitrary, and Τ G «£f (^, Ж). 77и?и
(2.5.1) p(7) + v(7)=dimr.
PROOF: Let {vp..., vz} be a basis for кег(Г), / = ν(Γ), and extend it
to a basis of ^ by adding {u{,..., uk}. By 1.3.5 we have 1+к = а\тУ.
The theorem follows if we show that k = p(T). We do this by showing
that {Γιΐρ.. -,Тик] is a basis for range(r).
Write any ν G Г as £?=1 β·ν· + £f=1 fe-n·. Then 7 ν = £f=1 Ь-Ги,.
(since Γν- = 0). This shows that {Tux,..., Tuk} spans range(r).
To prove that {Tul,..., Ги^} is also independent, we observe that
if Σ*=ι cjTuj = °>then τ{Σ)=ι CjUj) = 0, that is, Σ)=ι CjUj G кег(Г).
Since {vp..., vj is a basis for ker(T), we have T!j=\ c\u\ — T!j=\ ^jvj
for appropriate constants d·. But {vp..., vj U {uv ..., и^} is
independent, and it follows that с = 0 (and d- = 0) for all j. <
2.5. Kernel, range, nullity, and rank
53
The proof gives more than is claimed in the theorem. It shows that Τ
can be "factored" as a product of two maps:
(2.5.2) У^У/кег(Т), and У/кег(Т)^ТУ.
The first is the quotient map У —> У/ ker(T); vectors that are
congruent modulo кег(Г) have the same image under 7\ which means that
Τ assigns the same image to the elements of each coset, and different
images to elements of distinct cosets. Assigning to cosets the
common image of its elements defines the second map, У j ker(T) —> TV
which is both injective and surjective, i.e., an isomorphism. (This is
the Homomorphism Theorem of groups in our context.)
2.5.2 The identity operator, defined by Iv = v, is an identity element
in the algebra Jf(y). The invertible elements in Jf (У) are the
automorphisms of У, that is, the bijective linear maps.
In the context of operators on finite-dimensional spaces, and more
generally, between vector spaces of the same finite dimension, injectiv-
ity and surjectivity are equivalent properties.
Theorem. Let Ψ be α finite-dimensional vector space, Τ £ Л£{У).
Then
(2.5.3) ker(7) = {0} 4=^ range(r) = Г,
and both conditions are equivalent to Τ being invertible.
PROOF: ker(7) = {0} is equivalent to ν(Γ) =0, and range(r) = ^
is equivalent to p(T) = агтУ. Now apply (2.5.1). <
It follows that for an operator Τ on a finite-dimensional space V,
invertibility and nonsingularity are equivalent.
2.5.3 As another illustration of how the "rank and nullity" theorem
can be used, consider the following statement (which can be seen
directly as a consequence of exercise exl.3.10).
Theorem. Let У = У{ 0 У2 be finite-dimensional, and άϊναΨχ = к.
Let W С У be a subspace of dimension I > k. Then dimW ПУ2 >
l-к.
54
2. Linear Operators and Matrices
PROOF: Denote by πλ the restriction to Ψ of the projection of Ψ on
Ух along У2- Since the rank of πχ is clearly < /c, its nullity is > / — k.
In other words, the dimension of the kernel of this map, dim W Π У2>
is at least l — k. <
EXERCISES FOR SECTION 2.5
ex2.5.1 Assume that У is finite-dimensional and Τ e а?(У,Ж). Let ν and
w be bases for У and W respectively. Prove that p(T) is equal to the rank
P(Ar.v.w) of the matrix Arvw.
ex2.5.2 Assume T,S e а?(У). Prove that v(ST) < v(S) + v(T).
ex2.5.3 Let Τ e &{Ψ\ and let Ψ С У be a subspace such that Ψ э кег(Г),
and ГЖ = ГГ. Prove that Ψ = У.
ех2.5.4 Give an example of two 2x2 matrices A and В such that p(AB) = 1
andp(£A)=0.
ex2.5.5 Assume dim У = n,T,Se ^(У). Prove
(2.5.4) р(Г5)=р(5)-dim (5Г П кег(Г)).
ex2.5.6 Given vector spaces У зала W over the same field. Let {v-}n-x d Ψ
and {wj}nj=l С W. Prove that there exists a linear map Τ of spanfvj,..., v„]
into W such that 7V, = wj for all 7 if and only if the following implication
holds:
η η
If a , j = 1,..., и, are scalars, and Σαΐνΐ = ^' ^еп Ha/W/ = ^-
1 1
Can the definition of Τ be extended to the entire У ?
ex2.5.7 What is the relationship of the previous exercise to Theorem 1.4.5?
ex2.5.8 The operators T,S £ а?(У) are called "equivalent" if there exist in-
vertible А,Б G if (У) such that
S = ATB (so that Τ =A~lSB-]).
Prove that if У is finite-dimensional, then Г,£ are "equivalent" if and only if
P(S)=P(7-).
2.5. Kernel, range, nullity, and rank
55
ex2.5.9 Give an example of two operators on R3 that are "equivalent" but
not similar.
ex2.5.10 Assume Г,5Е Л?(У). Prove that the following statements are
equivalent:
a. ker(S) С кег(Г);
b. There exists R e а?(У) such that Τ = RS.
Hint: For the implication a => b: choose a basis {vj,..., vs} for ker(S).
Expand it to a basis for кег(Г) by adding {ux,..., ut_s}, зала expand further
to a basis for У by adding the vectors {wx,..., wn_t}.
The sequence {Sul,... ,Sut_s} U {Sw^,... ,Swn_t} is independent, so that
R can be defined arbitrarily on it (and extended by linearity to an operator on
the entire space). Define R(SuA = 0, R(Swj) = Tw ·.
The other implication is obvious.
ex2.5.11 Assume Г.5Е Л?(У). Prove that the following statements are
equivalent:
a. range(S) С range(T);
b. There exists R e аС(У) such that S = TR.
Hint: Again, b => a is obvious.
For a=> b take a basis {vl5...,v„} for У. Let и-, j= Ι,.,.,η, be such
that Tu- = Sv: (use assumption a). Define Rvj = u· (and extend by linearity).
ex2.5.12 Find bases for the nullspace, ker(A), and for the range, range (A),
of the matrix, acting by right-multiplication on (rows in) R~!,
Γ1 0 0 5 9]
0 1 0 -3 2
0 0 1 2 1 .
3 2 1 11 32
Ь 2 0 -1 \Ъ\
ex2.5.13 Let Τ G jSf (V), / G N. Prove:
a. кег(Г;) С кег(Г/+1); equality if and only if range(rz) Π кег(Г) = {0}.
b. range(r/+1) С range(rz); equality if and only if кег(Г/+1) = ker(rz).
с If кег(Г/+1) = ker(rz), then ker(Tl+M) = ker(Tl+k) for every positive
integer к.
56
2. Linear Operators and Matrices
*ex2.5.14 Prove that the rank of a skew-symmetric matrix is even. (See 1.2.3
example d.)
*2.6 Operator norms
2.6.1 If Ψ and W are normed finite-dimensional vector spaces (see
* 1.5), we define a norm on jSf (Г, Ψ) by writing, for Τ e JSf (Г,:
(2.6.1) ||Г|| = тах||Гу||=тах-'
||v||=i v^o ||v||
Equivalently,
(2.6.2) ||Γ|| = inf{C : ||Γν|| < C||v|| for all ν e Г}.
To check that (2.6.1) defines a norm, we observe that properties a and
b (in 1.5.1) are obvious, and that с follows from
\\(T + S)v\\ < \\Tv\\ + \\Sv\\ < ||Γ||||ν|| + ||5||||ν|| < (imi + ||S||)||v||.
The norms appearing in the inequalities are the ones defined on W,
Jf(y,W)9 and У, respectively.
The space ££{Ψ) is an algebra and we observe that the norm
defined by (2.6.1) on &{У) is submultiplicative: given 5, Τ e 3?(У),
then, for every ν e У, \\STv\\ < ||5||||Γν|| < ||5||||Γ||||ν||, which means
(2.6.3) ||57|| < ||5||||Γ||.
Remark: If Ψ is real or complex, so is Jz?(У) and, as remarked in
1.5.1, the notions of point-set topology are well defined, independently
of the particular norm that we may be using.
EXERCISES FOR SECTION 2.6
ex2.6.1 Let У be a normed linear space and Τ £ ££(У). Prove that the set
of vectors vGf whose Γ-orbit, {Tnv}, is bounded is a subspace of Ψ.
Chapter 3
Duality of Vector Spaces
3.1 Linear functional
All the vector spaces we consider in this chapter are assumed to be
finite-dimensional.
Definition: A linear functional, a.k.a. linear form, on a vector space
У is a linear map v* of У into the underlying field F.
As mentioned already in 2.1.3, the space Jf(y,F) is called the
dual space of У, and is commonly denoted У*. As seen in 2.1.5, У*
with the standard operations (addition of operators and multiplication
of operators by scalars) is a vector space over F.
3.1.1 Linear functionals arise naturally in the expansion of vectors
as linear combinations of elements of a basis.
Let {vj,..., v„} be a basis for Ψ. Every element ν e У can be
written, in exactly one way, as
(3.1.1) ν = Σαλν>Γ
ι
The notation aAv) makes explicit the dependence of the coefficients
on the vector v.
Let ν = Li tf,(v)v and и — YJ\aj(u)v ,·· If c, d Ε F, then
η
cv-\-du — Σ(€αί(ν) +da:(u))Vj
ι
SOthat / J ^ / ч J / ч
aAcv + du) = caAv) +daAu).
In other words, the coefficient functions a -(v), j = 1,... ,n, are linear
functionals on Ψ.
57
58
3. Duality of Vector Spaces
By Theorem 2.1.2, linear functionals are completely determined
by the values they assign to the elements of any basis and, given a
basis, one can assign arbitrary values to its elements and extend by
linearity to a linear functional. The coefficient (or coordinate)
functionals clAv) of (3.1.1) can be defined in this manner, simply by the
fact that aAv) assigns the value 1 to ν · and 0 to vk for к ф j.
A standard notation for the image of a vector ν under a linear
functional v* is (v,v*). Accordingly we denote the linear functionals
corresponding to aAv) by v* and write aAv) = (v, v*) so that
(3.1.2) v = ][>,v;*)v;.
j
Proposition. The linear functionals v*, j = 1,..., n, form a basis for
the dual space У*.
Proof: Spanning. Let w* e У*\ write bj(u*) = (v -,и*). We claim
thatM*=Lfey.(ii*)vJ.
By (3.1.2), if ver,
OmO = (I(v,v*)v;.,W*) =£(v,vj)fty.(n*) = (ν,Σ>>>3),
j J J
and, since assigning the same value to every ν еУ means equality for
functionals, we have u* = L^,(w*)v*.
Independence. For any coefficients c· gF, j = 1,...,w, we
have ck = (νΛ, £c -vp, so that if £c;v* = 0, then ck = 0 for all к. М
Corollary. dim Ψ * = dim Ψ.
The basis {v*}", j= Ι,.,.,η, is called the dual basis of {v 1?... ,v„}.
It is characterized by the condition
(3.1.3) (v ν.) = δ /ι ^ = к;
0 otherwise.
δ ·, is known as the Kronecker delta.
3.1. Linear functional
59
3.1.2 The vector space structure defined on У* guarantees that for
every ν e Ψ the map ν* η^ (ν, ν*) is a linear map from Ψ* to F, that is,
a linear functional on Ψ*.
If {vv...,vn} is a basis for Ψ, and {vj,...,v*} the dual basis in
Ψ\ then (3.1.3) identifies Г as the dual of Г* and {vp...,vn} as
the dual basis of {v\,..., v*}. This means that every linear functional
on Ψ* is obtained as v* н^ (ν,ν*) for some ν e У. The roles of Ψ
and У are perfectly symmetric, and what we have is two spaces in
duality, the duality between them defined by the bilinear (that is, linear
in ν for every fixed v*, and linear in v* for every fixed v) form (v, v*).
Expansion (3.1.2) works in both directions; thus if {vj,...,v„} and
{vj,..., v*} are dual bases, then for all ν e Ψ and v* e У*,
(3.1.4) v = £(v,vj)vy., v* = £(v;,v>*.
1 1
The dual of F£ (i.e., ¥n written as columns) can be identified with
F£ (i.e., ¥n written as rows) and the pairing (v, v*) as the matrix product
v*v of the row v* by the column ν (exercise ex3.1.5 below). The dual
of the standard basis of F£ is the standard basis FJ.
3.1.3 Annihilator. Given a set А С У, the annihilator of A is the set
AL of all the linear functionals v* G У that vanish identically on A.
Formally:
(3.1.5) А^КеГ*: VvGA, (v,v*)=0.}
Clearly, A1- is a subspace of У*.
Linear functionals that annihilate A vanish on span [A] as well, and
functionals that annihilate span [A] clearly vanish on its subset A; hence
A± = (span[A])±.
In particular, if {vj,...,vw} is a basis for a subspace Ψχ С Ψ, then
ν* e V if and only if (vy-, v*) = 0 for ; = 1,... ,m.
Theorem. Let ΨχαΨ bea subspace; then dim Ψχ + dim Ψ^ = dim Ψ.
60
3. Duality of Vector Spaces
PROOF: Let {v],...,vw}beabasisfor>/1, and let {vm+1,...,v„}
complete it to a basis for У. Let \y\,..., v*} be the dual basis.
We claim that {v^+1,..., v*} is a basis for У^. This will imply
that ахтУ^ — n — rn, proving the proposition.
By (3.1.3) we have {ν^+1,...,ν*} С У^~ and, being part of a basis,
these vectors are independent. We only need to prove that they span
yxL.
Let w* Ε Ψγ~, write w* = Ц-=1 α;·ν*, and observe that a- = (v ·, w*).
If w* Ε У^, then a- = 0for 1 < j < m, so that w* = £^+1 α -vj. ^
Corollary. Lei Ух С У be a subspace, and ν e У\УХ. Then there
exist v* Ε ^ swc/г that (v, v*) 7^ 0.
PROOF: As span[v,yj properly contains ^, its dimension is bigger
than that of Ух. Its annihilator has lower dimension than Ψ^~, and is
therefore properly contained in the latter. This is precisely our claim.
Remark: The use of dimension to prove the corollary is somewhat
heavy-handed and the proof is less transparent than the following
observation. In the notation of the proof of the theorem, the assumption
that ν <£УХ implies directly, by (3.1.4), that (ν, ν£) Φ 0 for some k> m
(so that v* Ε У^).
Restating the corollary for the case that Ух is given as spa η [A] for
some subset А с У, we have
Proposition. Let А С У, ν Ε У. Then ν Ε spa η [A] if and only if
(ν, и*) = 0 for every и* Ε Α1-.
3.1.4 Let У be a finite-dimensional vector space and Ух С У a sub-
space. The restriction of a linear functional in У to Ух is a linear
functional on Ух.
The functional whose restriction to Ух is zero are, by definition,
the elements of У^. The restrictions of v* and w* to Ух are equal if
and only if v* — w* Ε У^~. This, combined with exercise ex3.1.2 below,
gives a natural identification of Ух with the quotient space У*/У^.
3.1. Linear functionals
61
EXERCISES FOR SECTION 3.1
ex3.1.1 Given a linearly independent {vx,.. .,vk} СУ and scalars {а^=х.
Prove that there exists v* G У such that (v ·,v*) = a- for 1 < j < k.
ex3.1.2 If Ух is a subspace of a finite-dimensional space У, then every linear
functional on Ух is the restriction to Ух of a linear functional on У.
ex3.1.3 Let Ух С У be a subspace. Prove that (identifying (У*)* with У)
(yx±)± = yx.
ex3.1.4 Let У be a finite-dimensional vector space, and tjcfa subspace.
Let {w£}£=1 С У* be linearly independent mod У^ (that is: if Y,cku\ G
^χ, then c^ = 0, for к = 1,..., r). Let {v*}J=] С У,х be independent. Prove
that {w|} U {ν*} is linearly independent in У*.
ex3.1.5 Show that every linear functional on F£ is given by some (ax,..., an)
in F? as
Xn
(fl,,...,fl„)
:ΣβΛ·
ехЗ.1.6 Let У and Ж be finite-dimensional vector spaces. Prove
a. For every ν еУ and w* G ^* the map
9V|W, : Γ^(Γν,νν*)
is a linear functional on ££{У', #^).
b. The map ν 0 w* н^ <py w* extends to an isomorphism of У 0 /^* onto the
dual space of JSf(r,^).
ex3.1.7 Let У be a complex vector space, {v*H=] С У*, and w* G У* such
that for all vGf,
|(v,w*)|< max |(v,v*)|.
Prove that w* G span[{v*H J.
ex3.1.8 Linear functionals on RN[x}:
a. Show that for every χ G R the map <ρχ defined by (Ρ, <ρχ) = />(*) is a linear
functional on RN [x].
b. If {jtj,...,jcw} are distinct and m < N + 1, then φχ. are linearly
independent.
Hint: Write />z(x) = cl Π .^z(* -*,·) with cz = ПО*/ - *,·) ~!.
62
3. Duality of Vector Spaces
c. Prove Lagrange's theorem: Given m distinct numbers {xl,... ,xm} and m
arbitrary numbers { α j,..., aw }, there exists a polynomial Ρ of degree m - 1
such that Ρ(x A = а ■ for all j.
d. For χ G К and / G N, / < TV, the map <pW defined by (/>, <pW) = />(')(*)
(the /'th derivative of Ρ at jc) is a nontrivial linear functional on МдгИ-
e. Let (Xjjj), 7 = 1,... ,Af +1, *y- G Μ, Zy- G Ν, be distinct pairs. Denote by
#(m) the number of such pairs with Ζ ■ > m.
Prove that a necessary condition for the functional φ\j to be indepen-
j
dent on RN[x] is:
(3.1.6) for every m < N, #(m)<N — m.
Hint: If / > &, and χ is arbitrary, then φ№ G (R^x])1.
f. Check that φ{, φ_ j, and φ^ are linearly dependent in the dual of R2 [χ],
hence (3.1.6) is not sufficient. Are <p1? <p_p and φ^ linearly dependent
in the dual of Щ [x] ?
ex3.1.9 Let £? = R2[jc,y,z], the space of quadratic polynomials in three
variables with real-valued coefficients, that is,
(3.1.7) &> = {^ajUxjykzl : ; > 0, t > 0, Ζ > 0; 7 + Jfc + Z < 2}.
Denote by dM the directional derivative at χ = у = ζ = 0 in the direction mgM3,
i.e., the map Ρ \-^ duP(0). Denote by d}lv the corresponding mixed double
derivative at the origin χ = у = ζ = 0.
Prove that for all w, ν G M3, dw and c^y are linear functional on £?.
Prove that if μ, ν, w are linearly independent in M3, the functional
dw> dv^ dw, du in duv, duw, ayy, dvw, dww
are linearly independent. Do they form a basis for ^*?
3.2 The adjoint
3.2.1 Let Τ G &{?,W) and w* G Ж* = jSf (5^,F). The
composition
(3.2.1) w*T: v^(7v,w*)
is a linear map from Ψ to the underlying field, i.e., a linear functional
v* on У.
3.2. The adjoint
63
For 7 fixed, the map 7*: w* н-> w*7 from Ж* into У* is clearly
linear. It is called the adjoint of 7.
The basic relationship between 7, 7*, and the bilinear forms (v, v*)
and (w, w*) is: for all ν G У and w* G Ψ*,
(3.2.2) (7v,w*) = (v,7*w*).
Notice that the bilinear form on the left-hand side is the bilinear form
on (W, W*), while on the right-hand side it is the one on (V, V*).
If we identify the dual space Г** of Ψ * with Г (see 3.1.2), and the
dual Ж** of Ж* with Ж, then (3.2.2) also identifies 7 as the adjoint
7** of 7*.
3.2.2 We have seen in 3.1.2 that if Ψ = F£, Ж = F™, both with
standard bases, then we can identify ^* with FJ! and Ж* with F™, and
the standard basis of F£ is the dual basis of the standard basis of F£.
lfA=AT
hi ··· *1И
ж\
is the matrix of an operator 7 G «Sf (F£, F|
... r,
with respect to the standard bases, then the operator 7 is given as left-
multiplication by A on F£, and the bilinear form (7v,w), for wGFJ1
and ν G F£, is just the matrix product w(Av). We have
(3.2.3) (7v, w) = w(Av) = wAv = (wA) v,
and it follows that 7*w = wA7. That means that the action of 7* on
the row vectors in FJ? is obtained as right-multiplication by the same
matrix A =AT.
If we want1 to have the matrix of 7* relative to the standard bases
in F£ and F™, acting on columns by left-multiplication, all we need to
do is transpose w and wA and obtain (see also example 2.4.8)
7 w —A w .
lrrhis will be the case when there is a natural way to identify the vector space with
its dual, for instance when we work with inner-product spaces. If the "identification"
is through a sesquilinear form, as is the case when F = C, then the matrix for the
adjoint is the complex conjugate of Λ ; see Chapter 6.
64
3. Duality of Vector Spaces
3.2.3 Proposition. Let Τ G Л?(У, W), and let Г* be its adjoint. Then
(3.2.4) range(7)J- = кег(Г) and range(7*)± = кег(Г).
PROOF: The condition w* G range(r)-L is equivalent to (Γν,νν*) = 0
for all ν e У. Since (7V,w*) = (v,T*w*), the condition is equivalent
to (ν, Γ*νν*) = 0 for all ν G У, which is equivalent to 7*w* = 0. This
proves the claim range(r)-L = ker(7*).
The second equality in (3.2.4) is the same statement as the first,
applied to Γ* and its adjoint 7** = Τ instead of Τ and its adjoint T*.
Corollary. Let Τ G Jz?(^', W) and let Γ* fee its adjoint; then
(3.2.5) p(T*)=p(T).
PROOF: Denote η = dim Г = dim Г * and m = dim Ж = dim>T*.
Denote alsop(r) = dimrange(T) andp(r*) = dimrange(r*).
By Theorem 3.1.3, dimrange(r)"1 = m — p(T). By Theorem 2.5.1,
dimker(7*) =m-p(T*). By (3.2.4) the two are equal. A
EXERCISES FOR SECTION 3.2
ex3.2.1 If У = Ψ 0 <% and 5 is the projection of Ψ on Ψ along <%
(see 2.1.1, example /), what is the adjoint 5*?
ex3.2.2 Show that Proposition 3.2.3 is equivalent to the statement that
the column rank of a matrix is equal to its row rank.
ex3.2.3 A vector ν G Ψ is an eigenvector for Τ G Л?{У) if Tv = A v
with A G F; A is the corresponding eigenvalue.
Let ν G У be an eigenvector of Г with corresponding eigenvalue A,
and w G f * an eigenvector of the adjoint Г* with eigenvalue A* 7^ Я.
Prove that (v,w*) = 0.
Chapter 4
Determinants
4.1 Permutations
As mentioned in example с of 1.1.1, a permutation of a set X is
a bijective map of X onto itself, and the collection S(X) of all the
permutations of X forms a group under the operation of composition of
maps: the product τσ of the permutations σ, τ Ε S(X) is defined by:
for j Ε Χ,
(τσ)(;) = τ(σ(;)).
The identity element of S(X) is the permutation e defined by e(j) = j
for all j eA (the trivial permutation).
We are interested mostly in permutations of the set X = [1,... ,n].
The group S„ = S([l,...,и]) is called the symmetric group on [1,..., n].
4.1.1 For σ Ε S„, the σ-orbit of an element β Ε [1,... ,n] is the set
{σ*(α)}. If β is a fixed point of σ, i.e., σα = β, its orbit is reduced to a
single point; we refer to such orbits as trivial.
A cycle is a permutation with a unique nontrivial orbit. Cycles σ
are often written as (αχ,..., ат), where {β ,·}7=ι is the unique nontrivial
orbit, enumerated in such a way that α ·+1 = σ(α ·) for I < j < m, and
«j = a(am).
Observe that σ is determined by the cyclic order of the entries,
thus
\fl\ j · · · ->am) z= \amia\ ·, · · · tam—\ )-
Every orbit Ο = {σ;β} of a permutation σ determines a cycle τ0 by
setting τ0 = σ on O, and τ0 = e, the identity, on the complement
[1,..., n] \ O. We refer to τ0 as the restriction of σ to O.
65
66
4. Determinants
Two cycles are said to be disjoint if their (nontrivial) orbits are
disjoint.
Lemma. Every σ £ Sn is a product ofpairwise disjoint cycles. The
representation as such a product is unique.
PROOF: The σ-orbits form a partition of [1,...,n], and σ is the
product of its restrictions to the various orbits. For the uniqueness observe
first that disjoint cycles commute; that means that the "uniqueness"
must allow for arbitrary reordering of the factors. The claim is that the
(unordered) set of factors is unique.
Now observe that if σ = Π^,·, and the factors τ are pairwise
disjoint cycles with orbits О ·, then the restriction of σ to О is equal to
τ since all the factors тк, к φ j, act like the identity on О . In other
words, writing σ as the product of its restrictions to the various orbits
is the only way to write it as a product of pairwise disjoint cycles. Μ
The σ-period of a point a is the number of points in its orbit; equiva-
lently, it is the first positive integer / such that σ\α) = a. Notice that
the order of a cycle σ = (a{,... ,am), that is, the smallest positive
integer / such that σι is the identity, is equal to its length m, that is, the
number of elements in its orbit.
4.1.2 Cycles of length 2 are called transpositions.
Lemma. Every permutation σ £ Sn is a product of transpositions.
PROOF: Since every σ £ S„ is a product of cycles, it suffices to show
that every cycle is a product of transpositions. Observe that
(av...,al) = (apava2,...,al_{) = (ava2)(a2,a3) · · · (al_val).
In words: al trades places with ax_v then with ax_v etc., until it settles
in place of a{; every other a- moves once, to the original place of a -+1.
Thus, every cycle of length / is a product of / — 1 transpositions. <
4.1. Permutations
67
*4.1.3 Conjugation in S„. If σ, τ G Sn, and τ (J) = j then τσ_1 maps
σ(/) to j and στσ~ι maps σ(/) to a(j). This means that the cycles
of στσ~ι are obtained from the cycles of τ by replacing the entries in
each (cycle of τ) by their σ-images. Conversely, if τ] and τ2 have the
same orbit structure, that is, have the same number of orbits of each
length, and σ is a permutation that maps every orbit of τ] on a τ2-
orbit of the same length keeping the order of the corresponding cycles,
we can verify, as above, that τ2 = σ^σ-1. This proves the following
proposition.
Proposition. Two permutations are conjugate in Sn if and only if they
have the same orbit structure. In particular, two cycles are conjugate
in Sn if and only if they have the same length.
4.1.4 The sign of a permutation. There are several equivalent ways
to define the sign of a permutation σ G S„. The sign, denoted sgn [σ],
is to take the values ±1, assign the value —1 to each transposition, and
be multiplicative: sgn [στ] = sgn [o]sgn [τ]; in other words, it is to be
a homomorphism of Sn onto the multiplicative group {1,-1}.
All these requirements imply that if σ can be written as a product
of к transpositions, then sgn [σ] = (— 1)*. But in order to use this as
the definition of sgn, one needs to prove that the numbers of factors in
all the representations of any σ G S„ as products of transpositions have
the same parity.
We introduce sgn in a different way:
Definition: We say that a set J of pairs {(&,/)} of distinct integers
from [1,... ,n] is appropriate for Sn if it contains exactly one of (j,i),
(/, j) for every pair /, j, 1 <i < j <n.
A simple example is J = {(ij) : 1 <i < j < n}. A more general
example of an appropriate set is: for τ G S„,
(4.1.1) Λ = {(τ(/),τ(;)):1</<;<η}.
If J is appropriate for S„, and σ G S„, then1
(4.1.2) Y\sgn(G(j)-o(i))= J] sgn(a(7)-^(/))sgn(7-i)
*<J {iJ)€J
The sign of integers has the usual meaning.
68
4. Determinants
since reversing the order of a pair (/, j) changes both sgn (o(j) — o(i))
and sgn (j — /), and does not affect their product. Notice that with
j = JT of (4.1.1) this takes the form
[|sgn (σ(;) - σ(ι)) = Y[sgn(στ(;) - στ(ί)) sgn(τ(;) - τ(ι)).
We define the sign of a permutation σ by
(4.1.3) Sgn[G}=Y\sgn(G(j)-G(i)).
i<j
Proposition. The map sgn : σ ι—> sgw [σ] is a homomorphism of Sn
onto the multiplicative group {1,-1}. The sign of any transposition is
-1.
PROOF: The multiplicativity is seen as follows:
sgn [στ] = Y[sgn (στ(;) - στ{ΐ))
= Y[sgn(oT(j)-OT(i))sgn(T(j)-T(i))Y[sgn(T(j)-T(i))
=sgn [σ]sgn [τ].
Since the sign of the identity permutation is +1, the
multiplicativity implies that conjugate permutations have the same sign. In
particular all transpositions have the same sign. The computation for (1,2)
is particularly simple:
sgn(;-l) = sgn(;-2) = l for all j>2, while sgn(l-2) = -l,
and it follows that the sign of every transposition is — 1. <
EXERCISES FOR SECTION 4.1
ex4.1.1 How many cycles of length η are there in S„? How many cycles of
length m are there in Sn?
ex4.1.2 Let σ £ S„, j = 1,2, be cycles with different orbits. Prove that the
two commute if and only if their (nontrivial) orbits are disjoint.
4.2. Multilinear maps
69
ex4.1.3 Let σ, τ be cycles of length / with the same nontrivial orbit. Prove
that the two commute if and only if each is a power of the other: a = xk and
τ = om. Can either к or m have a common factor with /?
Hint: km-I is divisible by /.
ex4.1.4 Let σ be a cycle of length k\ prove that sgn [σ] = (—\)k~l.
ex4.1.5 Let σ G Sn and assume that it has s orbits (including the trivial
orbits, i.e., fixed points). Prove that sgn [σ] = (—\)n~s
4.2 Multilinear maps
4.2.1 Let Ψρ j = 1,..., k, and W be vector spaces over a field F.
A map
(4.2.1) ψ: ijX^x-xf^f
is multilinear, or /c-linear (bilinear—if к = 2) if ψ(ν{,..., vk) is linear
in each entry ν · when the other entries are held fixed.
When all the У-'s are equal to some fixed У, we say that ψ is A:-
linear on У. If Ж is the underlying field F, we refer to ψ as a k-linear
form or just /c-form.
Examples:
a. Multiplication in an algebra, for example, (5,7") ι—> ST in «5?(У)
or (A,B) i-> AB in Ji{n).
К Let ^ = F[jc] and ^2 = F[y]; then the map (p(x),q(y)) ·-> p{x)q(y)
is a bilinear map from F[x] χ F[y] onto the space F[x,y] of
polynomials in two variables.
c. i//(v,v*) = (v,v*), the value of a linear functional v* G У* on a
vector ν £ ^, is a bilinear form on У χ ^*.
rf. Given к linear functionals v* G У*, their product
(4-2-2) ^)...,ν«(ν1,·..,ν,)=Π(ν;'ν})
is a &-form on У.
70
4. Determinants
4.2.2 If Ψ and Φ are ^-linear maps of Ψλ χ Ψ2 χ · ·. χ Ук into Ψ and
a,b e¥, then αΨ + ЬФ is /c-linear. It follows that the /c-linear maps
of ^ χ ^ χ · · · χ Ук into W form a vector space, which we denote
by J?3f({yj}kj=vW). When all the Ψ· are the same space Ψ, the
notation is: Jti£(У®\ W), and when Ψ = F, the reference to Ψ is
omitted. Thus, Jig(Уфк) is the space of all &-forms on Ψ.
4.2.3 Example d of 4.2.1 is very useful in that the linear span of the
forms it defines is the space of all &-forms on Ψ.
Theorem. Assume that Ψ is finite-dimensional. Let {u^,... ,un} be
a basis for Ψy and u* = {«*,... , w*} the dual basis. Let the forms
ψ« ν* be defined by (4.2.2). Then the set
1 '"■- к
(4.2.3) {ψν< vl:v*meu*, т=\,...,к}
1 '"■■' к
spans Jt^(Tek).
PROOF: We use induction on A:. If к = 1 there is nothing to prove.
Assume the result valid for к — 1. Let φ be a &-form. Write
(4.2.4) φ.(ν2,..., vk) = φ{μp v2,...,vk)
and observe that for every v] (as v] = Σ(ν\>и))ид>
(4.2.5) φ(ν,,ν2,...,νΛ) = £(ν1,Μρφ7.(ν2,...,νΛ).
The induction hypothesis applies to the A:— 1-forms φ , and expressing
φ in (4.2.5) as a linear combination of i//v* v* completes the proof.
J 1'···* fc-1
*4.2.4 The definition in 1.2.6 of the tensor product Ψλ®Ψ2 guarantees
that the map
(4.2.6) Ψ(ν,ιι) = ν®Μ
οΐΥλχΨ2 into Ψλ ® Ψ2 is bilinear. This map is special in that every
bilinear map from (Ψλ, У2) factors through it.
4.2. Multilinear maps
71
Theorem. Let φ be a bilinear map from (У\,У2) ^nto ^· Then there
is a linear map Φ: Ψχ®Ψ2 —> W such that φ = ΦΨ.
PROOF: The operator Φ is defined by: Ф(у®м) = Φ (ν, и). This is
unambiguous since for ν, G ^ and и ■ G ^,
£Vj®M.=o=*;[>(v.,M;.)=o.
One needs only to check that, so defined, the map Φ is linear, and that
φ = ΦΨ. We leave this to the reader. м
*4.2.5 Let У and Ж be finite-dimensional vector spaces.
If Τ G ЗС(У, Ж) is of rank 1, and w Φ 0 is in the range of T, we
have, for all ν G Ψ, Τ ν = a(v)w and the coefficient a{y) is a linear
form. In other words: there exists v* G Ψ* such that Tv = (v, v*)w.
On the other hand, for any choice of v* G У and w G Ж the map
0V^W defined by 0ν*Θνι;ν = (ν, v*)w is clearly a linear operator of rank
one from У ioW.
Theorem. The map Ф: v* ®w ^ ф^^ е £?(У ,W) extends by
linearity to an isomorphism ofY*®W onto ^£{Ψ, W).
PROOF: As in *4.2.4 we see that all the representations of zero in the
tensor product are mapped to 0, so that Φ can be extended
unambiguously to a linear map defined on all of У* ® W. Since the two spaces
have the same dimension, it is sufficient to show that Φ is surjective.
So let Τ G jSf (Г, W\ ν = {v;} a basis for Г, and v* = {v*} the
dual basis. Then, for ν еУ,
(4.2.7) Tv = г(Д>,у*)у,.) = Ι(ν,ν*)Γν. = (Σ>ν;βΓν;) ν,
so that Τ = ΣΦν*®Τν.- Th*s shows that Φ is surjective. 4
When there is no room for confusion, we simplify the notation and
write the operator as ν* ® w instead of 0V*^W.
72
4. Determinants
4.2.6 Symmetric and alternating A>forms. Let Ψ be a vector space
of dimension n, and let У* be its dual.
A bilinear form φ on Ψ is symmetric if (pfVpV^ = <p(v2,Vj) for
all Vj, v2 G Ψ\ it is alternating if <p(v, v) = 0 for all ν еУ.
Example: given a bilinear form ψ, define
(4 2 8) V^m(vi,v2)=y(v1,v2) + y(v2,v1),
Vre/i(vi,v2)=y(v1,v2)-y(v2,v1);
then the form i//5},w is symmetric, while i//fl/i is alternating.
More generally, a &-form φ is symmetric if
(4.2.9) Φ(νσ(1),..·,νσ(,)) = φ(ν1,...,ν,).
for every A: vectors v{,..., vk G У and every permutation σ G Sfc.
A /c-form φ is alternating, if φ(νι,...,ν1ζ) = 0 whenever two of the
entries are equal, i.e.,
(4.2.10) lfvj = vlforj^l, then φ(ν1,...,νΛ) = 0.
Condition (4.2.10) is equivalent to the seemingly stronger condition:
(4.2.11) Φ(νσ(1),...,νσ№)=^Λ[σ]φ(νρ...,ν^),
(for every к vectors Vj,..., vk G У and every permutation σ G Sk).
To show the equivalence we observe first that, since every
permutation is a product of transpositions, and since sgn [σ] is multiplicative
on Sk (see 4.1) it is enough to verify (4.2.11) for all transpositions
(ij). Using (4.2.10) and multilinearity, we have
φ(...,ν|.,...,ν7.,...) = φ(...,ν|.,...,ν;. + ν|.,...)
(4.2.12) =9(...,-vy.,...,vy. + vI.,...) = 9(...,-vy.,...,vI·,...)
= -φ(...,ν;.,...,ν|.,...),
proving (4.2.11) for transpositions, hence for all permutations.
The sets Jt^sym{1/®k) of all symmetric &-forms and Ji^alt(TQk)
of all alternating Axforms are linear subspaces of Ж^£(Уек).
4.2. Multilinear maps
73
EXERCISES FOR SECTION 4.2
ex4.2.1 Refine Theorem 4.2.3 and show that if {u\,..., w*} is a basis for ψ*,
then (w v* : v*m G {w^,..., w*},ra = 1,...,£} is a basis for <Ж£?(У®к).
1 '"■'' к
Hint: Ylkj=\ (uj , и* ) = 1 if jm = lm for all m and is zero otherwise.
ex4.2.2 Prove that the sets <Ж J*?sym(y®2) of symmetric bilinear forms on У
and ^J2?alt(y®2) of alternating ones are linear subspaces of JtJ£ (f®2),
and
(4.2.13) ^^(У®2) = ^т^(У®2) е^Г^ДГ®2).
ex4.2.3 If φ is a &-form on Ψ and (psym is defined by
(4.2.14) <Psym{vV...,Vk) = u Σ <P(vaW~va{k))>
' GeSk
then (psym is symmetric, and the map φ \-^ (psym is a projection of MJ£(У®к)
onto ^J2^,w(y®*). Similarly, the form cpalt defined by
(4.2.15) <Pfl/,(v1?...,V,) = - Σ ^W9(va(l)-'-va(*))'
is alternating, and the map φ \-^ φα1ί is a projection of ^#J^f (Уф/:) onto the
subspace^T^ir®*).
ex4.2.4 Let {^·} j < ·<η be a basis of У *. For ι < 7, let e. Л е- be the alternating
bilinear forms
eiAej(vvv2) = (vvei)(v2,ej)-(vvej)(v2,ei)
on У χ У. Prove that {е- Л ^}г<; is a basis for Jt^£alt
What are the dimensions of the spaces
^^(f®2), ^^m(f®2), and Jti£alt
ex4.2.5 With {ег}1<г<„ as above and for i < j < к, let ei А е- Аек be the
trilinear forms
(4.2.16) eiAejAek(vvv2,v3)= £ ^« [σ](νσ(1),^·)(νσ(2)^ ·)(νσ(3),^).
74
4. Determinants
Prove that each et A e- A ek is alternating and show that the set
{eiAejAek:i<j<k}
is a basis for М££г
alt
ex4.2.6 Assume that <p(v, и) is bilinear on Ψχ χ Ψ2·
Prove that the map Τ: и н^ <ρ(·, и) is a linear map from ^ into (the dual
space) Ψ{. Similarly, S: ν ^ <p(v, ·) is linear from Ψχ to Ψ2'·
ex4.2.7 Let Ψχ and ^ be finite-dimensional, with bases {v1?...,vm} and
{и15...,и„} respectively. Show that every bilinear form φ on (УХ,У2) *s
given by an m χ w matrix (a-k) such that if ν = Σ7·*7ν; anc* w = ТлУ^к tnen
(4.2.17) φ (ν, и) = ^Яд*,·^ = [χλ,... ,хт]
ι\η
У\
\Уп.
ех4.2.8 What is the relation between the matrix in ex4.2.7 and the maps S
and Τ defined in ex4.2.6?
ex4.2.9 Assume that Ύχ and Ύ2 are finite-dimensional, with bases \yx,..., vm}
and {ux,..., un} respectively. Let Τ e 3?(У{, Ψ2) and let
/iy —
*n\
be its matrix relative to the given bases. Let {v\,..., v^} be the dual basis of
{v1?...,vm}. Prove that
(4.2.18)
r = £a..(v*®w.)·
4.3 Alternating n-forms
4.3.1 If φ is an alternating η-form, and if one of the entry vectors in
φ(νχ,..., vn) is a linear combination of the others, we use the linearity
of φ in that entry and write φ{νχ,..., v„) as a linear combination of φ
evaluated on several η-tuples each of which has a repeated entry. Thus,
if {vx,..., vn} is linearly dependent, φ{νχ,..., v„) = 0. It follows that
if dim Ψ < η, there are no nontrivial alternating η-forms on У.
4.3. Alternating η-forms
75
Theorem. Assume that άιναΨ = η. The space of alternating n-forms
on Ψ is one-dimensional: up to multiplication by a scalar there exists
a unique nontrivial alternating η-form D on Ψ.
Moreover, D{yv... ,v„) Φ 0 if and only if {νλ,... ,v„} is a basis.
PROOF: Assume that dim У = n, letu= {ир...,и„} be a basis for У
and let {u\,...,w*} be the dual basis. The alternating n-form
(4.3.1) Du(vv...,vn)= £ «Λ[σ]Π(νσ(Λ,^)
aes„
is nontrivial since, as Π(Μσ( vMP = 1 when σ is the identity and
Π(^σ, ч,м*) = 0 otherwise, D{uv... ,и„) = 1.
It remains to show that if φ is an alternating η-form, then it is a
scalar multiple of Du. Specifically, we must show that
(4.3.2) φ(νι,..., vn) = q>{uv..., un)Du(vl,..., v„).
Let φ be an alternating η-form, then φ (и- ,...,w· ) = 0 if there is
a repeated index, that is, unless {^,.-.,Λ} is a permutation σ of
{l,...,w}, and then ф(ису(1),...,ису(я))=^л[а]ф(и1,...,ия).
If {vj,...,vn} is an arbitrary η-tuple, we express each vt in terms
of the basis {м15... ,и„}:
(4.3.3) ν; = Σαι,Λ' 7=1,·· ·,",
ι=1
where я- = (ν ·,«*) and the alternating multilinearity implies
9(v1,..-,v„)=£a1Ji...^J.^(W.i,...,W.J
(4.3.4) =(L sgn[a}alo{xy-ano{n^{uv...,un)
oesn
=0„(ν1,...,νΛ)φ(ιι1,...,ΜΛ).
This shows that the value of φ(^1,..., un) for a basis {i^,..., un}
determines <p(vj,..., vn) for all η-tuples. It also shows that all alternating
η-forms are proportional and that, unless φ is trivial, φ(νχ,..., ν„) Φ 0
for every independent set (i.e., basis) {vj,..., v„}.
76
4. Determinants
Notice that even though the coefficients at ■ in (4.3.3) depend on the
choice of the basis {ux,..., un] used, the value of the expression
Du(v1,...,v„)= £ sgn[G}alMiy-anMn)
GESn
depends only on the value of <p(uv... ,и„), where φ is an arbitrary
alternating n-form.
4.4 Determinant of an operator
4.4.1 Let D be a nontrivial alternating η-form on an n-dimensional
vector space Ψ, and let {vx,..., vn} be a basis for У.
Definition: The determinant detT of an operator Τ e Jf(y) is
(4.4.1) det7
D(Tvv...,Tvn)
D(vv...,vn)
The definition does pot depend on the choice of D; another choice is
just a nonzero constant multiple which appears in both the numerator
and the denominator, and hence does not affect the quotient.
The definition is also independent of the choice of the basis. If Τ is
singular, then {Tvx,..., Tvn] is linearly dependent (for any basis) and
the determinant is 0. If Τ is nonsingular and {w1?..., wn} is another
basis, express w, in terms of {v{,..., v„}:
(4.4.2) wj = T,ciJvp j=h--,n,
and observe that
(4.4.3) Twj = LciJTvp 7 = !>···>",
and the calculation in equation (4.3.4) shows that
D{Twv...,Twn) _ D(wv...,wn)
(4.4.4)
D{Tvv...Jvn) D(vv...,vn) '
which implies that using the basis {wp..., wn} would give the same
value as (4.4.1).
4.4. Determinant of an operator 77
IfA7v = (α· ·) is the matrix of Γ in terms of a basis ν = {vp... ,v„},
then
η
(4.4.5) Γν; = Σ^Λ' 7 = 1,···, л,
ι=1
and
(4.4.6) det7= Y,sgn[G]axMxy..anMn).
aesn
Notice that while the matrix AT v depends on the choice of v, the
determinant does not.
4.4.2 Proposition. Let Τ e &{V). Then det7 = 0 if and only if Τ is
singular.
PROOF: This was mentioned in the proof that detT is independent of
the basis used in its definition. Τ is singular if and only if it maps
a basis onto a linearly dependent set, and D(Tv],..., Tvn) = 0 if and
only if {Tvx,..., Tvn} is linearly dependent. Μ
4.4.3 Proposition. IfT,Se&{V) then
(4.4.7) det75 = det7det5.
PROOF: If either 5 or Τ is singular, then both sides of (4.4.7) are zero.
Otherwise detS φ 0, {Sv ·} is a basis, and by (4.4.1),
det7S = v , v ^ · -V1^ —^ = detTdetS.
D(Svv...,Svn) D(v,,...,v„) M
*4.4.4 Orientation. When Ψ is a real vector space, a nontrivial
alternating η-form D determines an equivalence relation among bases.
The bases {v ·} and {иЛ are declared equivalent if D(vx,..., v„) and
D(ux ,...,«„) have the same sign. Using — D instead of D reverses the
signs of all the readings, but maintains the equivalence. An orientation
on Ψ is a choice which of the two equivalence classes to call positive.
78 4. Determinants
4.4.5 Invariant subspaces. Let Τ G аС(У).
Definition: A subspace Fc^is Τ-invariant if w G Ψ implies
Tw eW.lfW is 7-invariant, the restriction Τψ, defined by: w i-> Tw
for w G W, is clearly a linear operator on Ψ.
Τ also induces an operator Τψ,ψ on the quotient space Ψ /W. The
operator Τγ,ψ is defined by
(4.4.8) Tr/w(v + TT) = Tv + W.
The coset v + W of a vector ν is mapped onto the coset of TV. The
definition is justified by showing that it is independent of the choice
of the representative: if νχ and v2 represent the same coset, that is,
Vj — v2 G Ж, then, as Ж is 7-invariant, 74^ — 7v2 = T(v{ — v2)^W
so that Tvx and Tv2 represent (belong to) the same coset.
The operator Τγ ,ψ is linear: write ν for the coset ν + Ψ\ then
Ty,w(alvl +a2v2) is the coset containing α(Τνχ + a2Tv2, which is the
sum for j =1,2 of thecosets aTv- + W, i.e., α{Τγ ιψνχ +a2Ty,wv2.
Proposition. IfW С Ψ is Τ-invariant, then
(4.4.9) det Τ = det Τψ det Τψ/ψ.
PROOF: Let {w ·}" be a basis for У, such that {wj}\ is a basis for Ψ.
If 7Ж is singular, then Τ is singular and both sides of (4.4.9) are zero.
If Ίψ is nonsingular, then w = {Twx,..., Twk} is a basis for W,
and {Tw^,..., Twk; wk+v..., w„} is a basis for Ψ.
Let D be a n on trivial alternating η-form on Ψ. Then
&(uv...,uk)=D(uv...,uk\wk+v...,wn)
is a nontrivial alternating k-foxm on W.
The value of D(Twx,..., Twk\ wfc+1,..., un) is left unchanged if the
variables ил+1, · · ·, un are replaced by ones that are congruent to them
mod W. In other words, denoting by й the equivalence class (mod W)
of vectors u, the form Ψ(δΛ+1,..., йп) = D(rwj,..., Гм^; ил+1,..., un)
4.5. Determinant of a matrix
79
is therefore a well-defined nontrivial alternating η — /c-form on У/Ж:
det Γ
D(Twv...,Twn)
D(wv...,wn)
D(Twx,..., Twk\wk+V ...,w„) D(7w1,..., Twn)
D(w1 ,...,w„) D(7w1,..., Twk\wk+l ,...,w„)
_Ф(7У1,...,7Ч) Ψ(7ν,+ 1,...,7ν„)
" ΦΚ,.,.,νν,) ' Ψ(*έ+1 вя) -det^detT^.
A special case of the proposition is seen in the following situation.
Assume that Τ G &(У), that Ψ = Ψχ θ Ύν and that both components
are T-invariant. One can then identify Ψ2 with Ψ/Ψχ and Τψ with
Ty,y, and the proposition reads
(4.4.10) det Τ = det ΤΨ det ΤΨ.
By induction on the number of direct summands we obtain the
following corollary:
Corollary. Assume that Ψ — 0 Ψ· and all the Ψ- 's are Ί-invariant.
Let Ίψ denote the restriction of Τ to Ψ^; then
(4.4.11) det7 = [|det7r.
EXERCISES FOR SECTION 4.4
ex4.4.1 Prove that if Γ is nonsingular, then detT-1 = (detT)-1.
4.5 Determinant of a matrix
4.5.1 Let A = (α·.) G M{n). The determinant of A can be defined
in several equivalent ways. Having defined the determinant of an
operator, we can define det A as the determinant of the operator TA that
A defines on ¥n by matrix multiplication. The standard definition is
80
4. Determinants
equivalent, done directly by the following formula, given by (4.4.6):
(4.5.1) detA
i2]
An\
= Y,sgn[G]aXM]y.a
oesn
η,σ(η)'
Having the option to use either definition is an advantage. The first
definition yields properties of detA from the properties of the
determinants of operators. For example: the fact that det(AZ?) = detA deti?
follows immediately from 4.4.3. On the other hand, the definition by
(4.5.1) is sometimes readier for computation.
4.5.2 Cofactors, expansions, and inverses. For a fixed pair (ij),
the elements in the sum in (4.5.1) that have a-■ as a factor are those for
which σ(/) = j. Their sum is
(4·5·2) Σ S8H [σΚ,σ(1) * * ·\σ(η) = aUAW
GES„.o(i)=j
The sum, with the factor ai}■ removed, denoted A- in (4.5.2), is called
the cofactorsi (i,j).
We leave the proof of the following lemma as an exercise.
Lemma. With the notation above, A· is equal to (—l)'+; times the
determinant of the {n — 1) χ (η — 1) submatrix obtained from A by
deleting the Vth row and the j'th column.
Partitioning the sum in (4.5.1) according to the value σ(/) for some
fixed index / gives the expansion of the determinant along its Vth row:
(4.5.3)
detA
ΣαυΑυ-
If we consider a "mismatched" sum: ]T . a^A^ for / φ к, we obtain the
determinant of the matrix obtained from A by replacing the /c'th row
by the /'th. Since this matrix has two identical rows, its determinant is
zero, that is
(4.5.4)
fori^k, LfliAy = 0·
4.5. Determinant of a matrix
81
Finally, write A
42
iA„
Кг
and observe that ]£. αί -ΑΛ ■ is
the /fc'th entry of the matrix AA so that equtions (4.5.3) and (4.5.4)
combined are equivalent to
(4.5.5) AA = detA/.
Proposition. If A G Ж{п) is nonsingular, thenA~x = dJu\A.
Historically, the matrix A was called the adjoint of A, but the term
adjoint is now used mostly in the context of duality.
EXERCISES FOR SECTION 4.5
ex4.5.1 Prove Lemma 4.5.2.
ex4.5.2 Let A G ^(/i,F). Prove that det(-A) = (-l)"detA.
ex4.5.3 A matrix A = {α· ·} G Jt{n) is upper triangular if α·. = 0 when ι > y.
A is /ои^ег triangular if α · · = 0 when i < j. Prove that if A is either upper or
lower triangular then detA = ΠΓ=ι αα·
ex4.5.4 The diagonal sum of the matrices A ·, j = 1,..., m is the matrix A that
has the m matrices Ax,..., Aw along the diagonal, and zero entries everywhere
else:
(4.5.6)
*i
0
0
0
A?
0
0
A,
0
0
0
0
0 0 0
Prove that detA = ndetA;.
\B С
ex4.5.5 Let A = , where A is an η χ η matrix, Б, С and D are
respectively mx m, mx (n — m), and (n — m)x(n — m) matrices. Give two
proofs that detA = detBdetD: one by applying Proposition 4.4.5, and the
second directly from definition (4.5.1).
82
4. Determinants
ex4.5.6 Prove: detA = det(A ). (A is the transpose of A: if A = (α· ·) then
A'= («„.).)
ex4.5.7 Let A G ^(n,R) be skew-symmetric. Prove that if η is odd, then
detA = 0.
ex4.5.8 Given a · G F, j = 0,..., η — 1, prove (compute the determinant) that
(4.5.7)
-Я
1
0
0
0
0
-Я
1
0
0
-Я
1
0
0
0
0
-Я
1
un-\
:(-1)И(ЯЯ+2>,.А>).
Η/'πί; Compute the cofactors of the elements of the last column.
ex4.5.9 How can the algorithm of reduction to row-echelon form be used to
compute determinants?
ex4.5.10 Let A G M(n). A defines an operator on F", as well as on M{n),
both by matrix multiplication. What is the relation between the values of
detA as operator in the two cases?
ex4.5.11 Prove the following properties of the trace:
1. If А,Б G Ji{n), thentrace(A+£) = traceA + traced.
2. If A G M(m,ri) and Б G <M(n,m), then traceAB = traceBA.
ex4.5.12 If А, Б G ^#(2), then (AB-BA)2 = -det(AB-BA)I.
ex4.5.13 Let A = (a- ■) G Jt(ri) and let m > n/2. Assume that ax■ ■ = 0
whenever both / < m and j < m. Prove that det(A) = 0.
ex4.5.14 The subset of <М{п,Ж) of matrices whose entries are integers
(called inegral matrices) is denoted Л((п,Ъ). Prove that an integral matrix
has an inverse in <M{n, Z) if and only if it is unimodular, that is, detA = ± 1.
ex4.5.15 The Vandermonde determinant. Given scalars а ·, j = 1,..., и,
the Vandermonde determinant V(a{,..., an) is defined by
V(av...,an) =
~n—\ |
1 ι
„n-1
1 aw
^и-1
4.5. Determinant of a matrix
83
Use the following steps to compute V(av...,an). Observe that
ν(αν...,αη,χ)
1 flj
1 flo
1 an a\
a
is a polynomial of degree η (in x).
a. Prove that V(ax,.. . ,α„,*) = V(al,...,«„)Π"=ι (*~я,·)·
b. Use induction to prove: V(al,... ,a„) = Пг<;(а; — ai)·
What is the rank of V(a1,...,α„)?
ex4.5.16 Let α , j = l,...,m, be distinct and let ^ be the space of all
trigonometric polynomials of the form P(x) = Σ!?=\ cijeiajx.
a. Prove that if Ρ £ g? has a zero of order ra (that is, a point x0 such that
pW (jc0) = 0 for / = 0,..., m — 1), then Ρ is identically zero.
b. For every k £ N there exist constants ck /9 / = 0,..., m — 1, such that if
PG^5 then />(*) (0) = Σ?^ο ск,1р{1) (°)·
c. Given {<:,},/ = 0,..., га- 1, there exists P G & such that pW (0) = cz
for 0 < / < ra.
Hint: P«(;c0) = £7=iaj(iaj)leiaJxo.
ex4.5.17 Let С G ^#(«,C) be nonsingular. Let 9tC, resp. 3C, be the matrix
whose entries are the real parts, resp. the imaginary parts, of the
corresponding entries in С Prove that for all but a finite number of values of a £ R, the
matrix 9tC + a3C is nonsingular.
Hint: P(x) = det(9tC + x3C) is a polynomial of degree < и, and P(i) φ 0.
ex4.5.18 Given that the matrices BvB2e М(п\Ж) are similar in ^#(«;C),
show that they are similar in Ж{п\Ж).
Hint: HCe^(n,C) is such that CBX = B2C, then Bv B2 G Л{п\Ж) implies
StC^ = £29tC and ЗС^ = B2 3C.
This page intentionally left blank
Chapter 5
Invariant Subspaces
The study of linear systems, that is, operators on a fixed vector
space Ψ, takes full advantage of the fact that Л?(У) is an algebra.
Polynomials in 7, that is, linear operators of the form P(T) = Σ я,· Т7',
where Ρ (χ) = Σα-χι Ε F[jc], play a crucial role in the understanding of
Τ itself. In particular they provide a way to decompose Ψ into a direct
sum of Γ-invariant subspaces (see 4.4.5 or below), on each of which
the behaviour of Τ is relatively simple. The key to this decomposition
is the minimal polynomial of T, and its prime-power factorization (see
A.6.3).
The underlying field, F, is not listed in the notation; it is assumed
known. This field may or may not be algebraically closed (see
Definition A.6.5). This property affects much of what is discussed in this
chapter, and it will appear as an explicit assumption in the statements
of results that depend on it.
5.1 The characteristic polynomial
5.1.1 The characteristic polynomial of an operator. Let (V,T)
be a linear system, and let {vj,..., vn} be a basis for ψ. Expanding the
expression D(Tv] — λνλ, ... , Tvn — Av„), which appears in the
definition of det (T — A), we see that the latter is a polynomial, in the
variable A, of degree η = άιν&Ψ and leading coefficient (—1)".
Definition: The characteristic polynomial'of an operator Τ e JSf (У)
is the polynomial %τ(λ) = det (T — A).
By Proposition 4.4.2, χτ(λ) = 0 if and only if Τ - A is singular,
that is, if and only if кег(Г — λ) φ {0}. The zeroes of χτ are called
eigenvalues of 7, and the set of eigenvalues of Τ is called the spectrum
85
86
5. Invariant Subspaces
of 7, and denoted σ(Τ).
For Я G o(T), the space кег(Г — Я) is called the eigenspace of Я.
Its nonzero elements (that is, the vectors ν φ 0 for which Τ ν = Я ν) are
the eigenvectors of Τ corresponding to the eigenvalue Я.
Theorem. 77ге characteristic polynomial is conjugation-invariant: if
S,T e Jf(y) are conjugate then χτ(λ) = χ8(λ).
PROOF: If S = RTR~l with R G GL(r) then, as det/Г1 =(det/?)"1,
det(S-A)=detfl(:T-A)/rl = det(r-A).
If we write #Γ = ЩаД·7', then ап — (—l)w5 and βο = detT". Each of
the coefficients a- of %T is conjugation invariant and in particular so
is the trace of Τ\ defined by tracer = (—\)η~λαη_λ.
5.1.2 The characteristic polynomial of a matrix.
Definition: The characteristic polynomial of a matrix A e Ж{п) is
the polynomial #л (Я) = det (A — Я).
The reader should verify that χΑ = χτ where 7^ is the operator
of left multiplication by A on F". The spectrum of the matrix A is,
by definition, o(TA), the set of zeros of χΑ = χτ . The following
proposition and its proof are essentially repetitions of Theorem 5.1.1.
Proposition. If А, В G <y#(n) яге similar then they have the same
characteristic polynomial. In other words, %A is similarity invariant.
PROOF: If В = CAC~\ then В - Я = С (А- Х)С~\ which implies
that
det(£^)=detCdet(A^) det(C_1) = det (А -Я). <
Remark: χΑ is not a complete invariant of the similarity class of A.
Matrices (or operators) that have the same characteristic polynomials
need not be similar. See exercise ex5.1.4 below.
5.1. The characteristic polynomial
87
5.1.3 Traces. The trace of a matrix A is defined as follows: write the
characteristic polynomial χΑ = Щя^А·7', then
(5.1.1) ап = {-\)\ a0 = detA, and αη_λ = (-Ι)""1 ][>,.
ι
The trace of A is trace(A) = (-1)η~]αη_ι = £?α...
The trace of an operator 7 is defined similarly:
trace(r) = (—l)n~]an_l where an_x is the coefficient of λη~Χ in #7.
It is the sum of the diagonal elements in any matrix that represents Τ
(relative to any basis).
Like any part of χΑ, the trace is similarity invariant.
*5.1.4 The entire characteristic polynomial can be recovered from the
set{trace(A*)}£lJ.
Proposition. For A £ Ж{п\Щу the coefficients of the characteristic
polynomial of A are polynomials in {trace(Afc)}^~J, with coefficients
in¥.
PROOF: Let Ap..., A„ be the eigenvalues of A (that is, the zeros of
χΑ) repeated according to their multiplicity, and lying perhaps in some
field extension of F, see A.6.4. Then
trace(A*) = £A* = ^(V--A),
and the proposition follows from Corollary A.6.8 in the appendix. <
EXERCISES FOR SECTION 5.1
ex5.1.1 If Ψ С Г is ^-invariant, then χτ(λ) = %T %T .
ex5.1.2 Let Τ £ Jf(Y) and let {νΜ=ι be eigenvectors of Τ corresponding
to distinct eigenvalues {λΜ=ι. Prove that the set {νΜ=ι is linearly
independent.
Hint: Observe that if Tv = Av, then Tlv = λ1 ν for all / £ N.
ex5.1.3 Let Τ £ ££(Ψ) and assume that o(T) consists of η = dim Ψ distinct
points. Prove that χτ (Τ) = 0.
88
5. Invariant Subspaces
ex5.1.4 Prove: the characteristic polynomial of an upper triangular η χ η
matrix A = (α( ·) is equal to Π"=ι (ai i~^)·
Let A and В be upper triangular nxn matrices with the same diagonal
elements, i.e., at · = b{ ·. Are the two necessarily similar?
ex5.1.5 Prove: the characteristic polynomial of the nxn matrix A = (at ·) is
equal to Π?=ι (ai i~ ^) plus a polynomial of degree bounded by η — 2.
ex5.1.6 Assuming F = C, prove that trace [a · ) is equal to the sum
(including multiplicity) of the zeros of the characteristic polynomial of (at: ·). In
other words, if the characteristic polynomial of the matrix [ai ) is equal to
Π"=ι (Я-А;), then £A. = Ια...
5.2 Invariant subspaces
5.2.1 Let (Ψ, Τ) be a linear system.
Definition: A subspace Ψχ с ^ is Τ-Invariant if Г^ С ^. The
entire space У and the trivial subspace {0} are T-invariant for every
Τ £jf(y). These are the trivial invariant subspaces.
If Ψχ is 7-invariant and vGfj, then TJν G Ψχ for all j, and
taking linear combinations of these we obtain that P(T)v G Vx for every
polynomial P. Thus, ^ is P(T)-invariant for every Ρ G ¥[x].
Remarks:
a. Both кег(Г) and range(r) are (clearly) T-invariant.
b. If 5, Τ G Jf(y) and the two commute, then ker(S) and range(S)
are both T-invariant. This can be seen as follows: S(Tv) = T(Sv) = 0
ifSv = 0.
For the 7-invariance of range(S) observe that TSY = S(TY) с
SY.
In particular, ker (P(T)) and range(P(T)) are T-invariant for every
polynomial Ρ in ¥[x].
с Given ν G У, the set spa η [7, v] = {P(T)v : Ρ G ¥[x}} is clearly
a subspace, clearly ^-invariant, and clearly the smallest T-invariant
subspace containing v.
5.2. Invariant subspaces
89
5.2.2 Assume again that 7,S G &(У)9 and TS = ST, then:
a. T commutes with P(S) for every polynomial P\ hence ker(P(5))
and range(P(S)) are 7-invariant (see 5.2.1 b). In particular, for every
Я G F, ker(5 — Я) is T-invariant.
b. If W is an 5-invariant subspace, then TW is S-invariant. This
follows from
STW = TSW С ΤΨ.
There is no claim that W is Г-invariant (an obvious example is 5 = /,
which commutes with every operator 7\ and for which all subspaces
are invariant). Thus, kernels offer "a special situation".
с If ν is an eigenvector for 5 with corresponding eigenvalue Я, i.e.,
ν G ker(5 — Я) (see 5.1.1), and if (the Γ-invariant subspace) ker(5 — Я)
is one-dimensional, then ν is an eigenvector for 7.
If dimker(5 — Я) > 1, Τ maps ν onto a vector in ker(5 —Я), which
may or may not be a scalar multiple of v. Consider the example S = 1:
S commutes with every Τ G «5f(^), ker(5- 1) = У, so that every
vector is an eigenvector of 5, and this gives no information about its
image under an arbitrary T.
5.2.3 We note that each eigenvector of Τ spans a one-dimensional
7-invariant subspace.
Recall that the spectrum of 7\ σ(Γ), is the set of all the
eigenvalues of T. It is the set of zeros of the characteristic polynomial of 7\
χτ(λ) = ά&(Τ-λ) (see 5.1.1).
If the underlying field F is algebraically closed every nonconstant
polynomial has zeros in F and hence the spectrum of every operator
Τ G Jf(y) is nonempty.
Proposition (Spectral mapping theorem). Let (У,Т) be α linear
system, λ G o(T), and Ρ G F[jc]. Then
a. P{X) G σ(Ρ(Τ)).
b. If F is algebraically closed, then
σ(Ρ(Τ)) = {Ρ(λ):λβσ(Τ)}=Ρ(σ(Τ)).
90 5. Invariant Subspaces
Proof: a. Let νλ be an eigenvector for Я, i.e., Τνλ = λνχ. Then
Thx = Vvx and Ρ(Τ)νλ = Ρ(λ)νλ.
b. Assume that F is algebraically closed. For μ G F, denote by с Α μ)
the roots of P(x) — μ, and by m their multiplicities, so that
Ρ{χ)-μ = Υ[{χ-α]{μ)Τι, and Ρ(Γ) -μ = Ц(Т-c^))mJ.
Unless с j (μ) G ο (Τ) for some 7, all the factors are invertible, and
hence so is their product. Μ
Remark: If F is not algebraically closed, σ(Ρ(Τ)) may be strictly
larger than Ρ(σ(Τ)). For example, if F = R, Τ the rotation by π/2 on
M2, and P(x) = x2, then σ(Τ) = 0 while σ(Τ2) = {-1}.
*5.2.4 Part a of the proposition can be refined as follows:
Proposition. Let (V,T) be a linear system, Я G o{T), and Ρ G ¥[x].
Then for all keN, ker((P(7) - Ρ(λ))*) D ker((7 - λ)*).
PROOF: P{x) — Ρ (λ) vanishes for χ = Я, and hence it is divisible by
χ-λ (see A.6.1). If Q G ¥[x\ and P(x)-Ρ(λ) = β(*)(*-λ), then
(P(*) -Ρ(λ))* = β*(*-λ)*, and (P(T)-P(X))k = Qk(T)(T-X)k.
Hence, if (Γ-λ)*ν = 0, then (Ρ(Γ)-Ρ(λ))^ = β*(Γ)(Γ-λ)*ν = 0.
5.2.5 As mentioned above, Γ-invariant subspaces are Ρ (Τ)-invariant
for all polynomials P. The converse, however, is not necessarily true.
A subspace W can be T2-invariant and not be Γ-invariant.
Examples on R2: 7\ the rotation by π/2, has no nontrivial invariant
subspace while every subspace is invariant under its square T2 = —/.
Similarly, the reflection 5 which maps (x,y) to (y,x) has the diagonal
{(χ,χ) : χ G Щ as the only 5-invariant subspace and yet S2 = /, the
identity, and "everything" is 52-invariant.
5.2.6 Theorem. Assume that Ψ is finite-dimensional and F
algebraically closed. Then:
a. Every Τ G Α£{Ψ) has one, or more, eigenvectors.
5.2. Invariant subspaces
91
b. IfS and Τ commute, then they have a common eigenvector.
с If J2, С <5£{У) is a set ofpairwise commuting operators, then there
is a vector ν which is an eigenvector of every Τ e J2.
PROOF: a. This is an immediate consequence of the fact that cf(T)
is nonempty. If Я G <т(Г) then ker(7 — Я) is a nonempty invariant
subspace, and every vector in it is an eigenvector for T.
b. If Я G o{T), we have seen that ker(7 — Я) is S-invariant.
By part a there is a vector ν G ker(7 — Я) (hence, an eigenvector
of T) that is an eigenvector for Sker(T_x\, and hence for 5.
с If Ψ is one-dimensional, there is nothing to prove: every ν G Ψ is
an eigenvector for every Τ G «Sf (У). So assume άιχηΨ > 1. Let srf be
the set of all nontrivial subspaces of Ψ that are 7-invariant for every
Claim: stf is nonempty. If all the operators in В are scalars (scalar
multiples of the identity), then every proper subspace is in gf.
Otherwise take a non-scalar TY e J2, λ{ e θ'(71); then Ψχ = ker(7j —λλ) is
a nontrivial subspace that is 7-invariant for every Τ G =2.
Let Ж be a minimal element in srf (see exl.3.8); we claim that
W is one-dimensional. Otherwise, the argument just given that srf
is nonempty applies to subspaces of W and we obtain a nontrivial
subspace WxdW that is Γ-invariant for all Τ G J2.
Every nonzero vector w Ε Ж is a common eigenvector of all the
operators in «=2. <
5.2.7 Theorem. Let W С У be a subspace, and Τ G ££{?). The
following statements are equivalent:
a. W is Τ-invariant;
b. W1- is T*-invariant.
PROOF: Remember that for all w G У and w* G У we have
(rw,w*) = (w,rV).
Statement a is equivalent to:
92
5. Invariant Subspaces
"for all w G W and w* G Ж-1 the left-hand side is identically zero."
Statement b is equivalent to:
"for all w G Ψ and w* G Ж-1 the right-hand side is identically zero."
5.2.8 The fact that, when F is algebraically closed, every Те&(У)
has eigenvectors, applies equally to the adjoint system (У*,Т*).
Let У be η-dimensional and let w* G У* be an eigenvector for 7*;
then Ψη_λ = [w*]-1 = {v G ^ : (v,и*) = 0} is a ^-invariant subspace of
dimension n— 1.
Repeating the argument in >^_1 we find a Γ-invariant ^_2 С Vn_\
of dimension η — 2. Repeating the argument a total of η — 1 times we
obtain:
Theorem. Asswme ito F is algebraically closed, and let Ψ be an n-
dimensional vector space over F. For every Τ G Jf(Y), there exists a
complete flag {УЛ, j = 0,..., n, of Τ -invariant subspaces of Ψ. That
means:
^o = {°}> % = V\ ^.iCK;, and dim^. = ;\
Corollary. If the field F is algebraically closed, then every matrix
A G Ж{п\Щ is similar to an upper triangular matrix.
PROOF: Apply the theorem to the operator Τ of left multiplication by
A on F£. For every j G [1,...,n], choose Vj in ψ. \ ψ._λ.
For / < n, the set {vp ..., vz} is a basis for the 7-invariant Ψχ so
that Ίν{ is a linear combination of {vp ..., vj, and the matrix В
corresponding to Τ in the basis {vj,..., vn} is (upper) triangular.
The matrices A and В are similar since they represent the same
operator relative to two bases. <
Observe (again) that the assumption that F is algebraically closed
is essential. If the underlying field is R (which is not algebraically
closed) and Г is a rotation by π/2 on R2, Τ admits no nontrivial
invariant subspaces.
5.3. The minimal polynomial
93
EXERCISES FOR SECTION 5.2
ex5.2.1 Let Ψ С Ύ be ^-invariant, and Ρ a polynomial. Prove:
a. P{T)W=P{TW).
b. P\l )ψιψ — Ρ\*-ψιψ)·
ex5.2.2 Let Ψ be ^-invariant. Prove that ker(7^) = кег(Г) Π Ψ.
ex5.2.3 Prove that every upper triangular matrix is similar to a lower
triangular one (and vice versa).
ex5.2.4 If ψχ С Ψ is a subspace, then the set {S : S <E -Sf (У), S^ С Уг} is
asubalgebraof^(r).
ex5.2.5 Show that if S and Г commute and ν is an eigenvector for S, it need
not be an eigenvector for Τ (so that the assumption in the final remark of
5.2.5 that ker(S — A) is one-dimensional is crucial).
ex5.2.6 Prove Theorem 5.2.8 without using duality.
Hint: Start with an eigenvector ux of T. Set ^ = spanfwj.
Let й2 G У/^{ be an eigenvector of Ty,^ , u2 G У a representative of
й2, and ^2 = span[M1?M2]. Verify that ύϊ/2 is ^-invariant. Let w3 G Ψ j^/2 be
an eigenvector of Τψ/^ , etc.
5.3 The minimal polynomial
5.3.1 The minimal polynomial for (Γ,ν). Assume now that У
is an η-dimensional space. Given Τ G <$?(Ψ) and vGf, let m be
the first positive integer such that {Th}™ is linearly dependent or,
equivalently, that Tmv is a linear combination of {Т^у}™~\ e.g.,1
m—1
(5.3.1) 7wv = - J^ajTjv.
о
Notice that the assumption that {Th}™~x is independent guarantees
that m < η and that the coefficients a- are uniquely determined.
Definition: The polynomial minPT y(x) = χ?" + £™_1 я .χ7', with я.
defined by (5.3.1) is called the minimal polynomial for (Γ, v).
lrThe minus sign is there to give the common notation: minPTv(x) = xm +
94
5. Invariant Subspaces
Theorem. minPT v (x) is the monic polynomial Ρ of lowest degree that
satisfies P{T)v = 0.
PROOF: {Tj'v}™~1 is independent. <
Remark: The set 9Tr v = {P e ¥[x] : P(7)v = 0} is an ideal in ¥[x],
(see A.6.1). The theorem identifies minPTv as its generator. In other
words: given ν e У and Ρ e ¥[x], then P(T)v = 0 if and only if Ρ is
divisible by minPT v
As a simple consequence of this remark, we obtain the following
important proposition.
Proposition. For Ρ e ¥[x], P(T) = 0 if and only if Ρ is divisible by
minPT v for every vGf.
PROOF: P(T) = 0 if and only if P(T)v = 0 for every vGl <
For к > 0 we have Tm+kv = [J'^T^v, and induction on к
proves that Tm+kv e span[v,..., Tm~] v]. It follows that {Tjv}™~x is a
basis for spa η [7, ν], and dim spa η [7, ν] = degminPTv.
If Ρ is a polynomial, wG^an arbitrary vector, and P{T)u = 0
then P(T)Tku = TkP(T)u = 0 for all к e N so that P(T) is zero on
spa η [7, и]. In particular minPTv(7) = 0 on spa η [7, v].
5.3.2 Cyclic vectors. A vector ν e У is cyc//c for the system (У,Т)
if span[7,v] = ^. Equivalently, ν is cyclic for (У,Т) if minPT v is a
polynomial of degree n. Not every linear system admits cyclic vectors;
consider 7 = /; systems that do are called cyclic systems.
If ν is a cyclic vector for (У, 7) and minPT y(x) = xn + ^_1 a-x^
then the matrix of 7 with respect to the basis ν = {ν, 7ν,..., Τη~ι ν}
has the form
ГО 0 0
10 0
(5.3.2) Ат у =01 0
[θ 0 1
-а0
-а2
5.3. The minimal polynomial
95
The characteristic polynomial of Τ is equal to det(A7 v — λ I) which
was shown in exercise ex4.5.8 to be equal to (—\)η(λη + ΣαД;).
Hence
(5.3.3) *r(A) = (-iyminPTv(A).
This implies, in particular, that if the operator Τ has a cyclic vector,
then XT(T) = 0. It is a special case, and a step in the proof, of the
following theorem.
Theorem (Cayley-Hamilton). χτ(Τ) = 0.
PROOF: We will show that χτ is a multiple of minPT u for every и G
У, and the theorem will follow from Proposition 5.3.1.
LetwG^, У/ = span[7,w],andminPTu = Aw + ^p1<^. The
vectors u, Tu,..., Tm~xи form a basis for ^. Complete {ТЫ}™~Х to a
basis for У by adding appropriate vectors wx,..., wn_m. Let AT be the
matrix of Τ with respect to this basis. The top left mxm submatrix of
AT is the matrix of T<%, and the (n — m)xm rectangle below it has only
zero entries. It follows that χτ = χΤ/ Q, where Q is the characteristic
polynomial of the (n — m)x(n — m) lower right submatrix of A.
By (5.3.3) applied to 7^, we have χτ^ = (—l)wminPTu, so that
Χτ = Xtv Q *s divisible by minPT u, as claimed. <
An alternate way to word the proof, and to prove an additional
claim along the way, is to proceed by induction on the dimension of
the space Ψ. The additional claim is:
Proposition. Every prime factor of %T divides minPTu for some
vector и^У.
PROOF: We reprove the theorem and prove the proposition.
a. If η = 1, both claims, of the theorem and of the proposition, are
obvious.
b. Assume the claims valid for all systems of dimension smaller than
n. Letuey,u^ 0, and ^ = spa η [7, и]. If У/ = У, the claims
are a consequence of (5.3.3) as explained above. Otherwise, both tf/
96
5. Invariant Subspaces
and Ψ1% have dimension smaller than η and, by Proposition 4.4.9
applied to Τ — λ (exercise ex5.1.1), we have χτ = χΤα χτ . By the
induction hypothesis, χζ (Ty,^) = 0, which means that χζ (Τ)
maps Ψ into ^, and since χτ^ (Τ) maps <$/ to 0, we have %T{T) = 0.
Each prime factor of χτ is either a factor of χτ or of χτ
and, by the induction hypothesis, is either a factor of minPTu or of
minPT λ~ for some ν = v+ <$/ G Ψl6^. In the latter case, observe
that minPTv(T)v = 0. Reducing mod ty/ gives minPT v(7^,^)v = 0,
which implies that minPT ~ divides minPT v. <
r ir/^,v i,v
5.3.3 We return to the matrix defined in (5.3.2). Given an arbitrary
monic polynomial, P(x) = jc" + £fc -χ7', the matrix
(5.3.4)
Ό 0
1 0
0 1
0 0
0 -Ъл
yn-\-
is called the companion matrix of the polynomial P.
If {w0,...,ww_1} is a basis for Ψ, and we define 5 G Л?(У) by
Su-
и ·+1 for j < η — 1, and Sun_
Lt^juj,
then w0 is cyclic for (^,5), the matrix (5.3.4) is the matrix ASu of 5
with respect to the basis u = {м0,...,м/1_1}, and minPSu = P.
Thus, every mon/'c polynomial of degree η is minPs u, the minimal
polynomial of some cyclic vector и in an η-dimensional system (Ψ,S).
5.3.4 The minimal polynomial. Let Τ G «5f(Г). The set 9Tr =
{P:Pe F[x], P(T) = 0} is an ideal in F[*]. The monic generator2 for
9ΐΓ is called the minimal polynomial of Τ and denoted minPT. To put
it simply: minPT is the monic polynomial Ρ of least degree such that
P(T) = 0. Similarly, the minimal polynomial minPA of a matrix A G
2SeeA.6.1.
5.3. The minimal polynomial
97
Ж{п, F) is the monic polynomial Ρ of least degree such that Ρ (A) = 0.
It is the minimal polynomial of the operator TA.
Since the dimension of Jz?(У) is n2, any n2 +1 powers of Τ are
linearly dependent. This proves that 9ΐΓ is nontrivial and that the degree
of minPT is at most n2. By the Cayley-Hamilton theorem, χτ G 9T7,
which means that minPT divides %T and its degree is therefore no
greater than n.
The operator Τ is called derogatory if the degree of minPT is smaller
than the dimension n, and nonderogatory if the degree is n, i.e., if
minPT = ±#Γ. Similarly, a matrix A is nonderogatory if minPA =
±#A, and derogatory otherwise, i.e., when deg^A > degminPA.
The condition "P(T) = 0" is equivalent to "P(T)v = 0 for all ν G
У, and the condition "P(T)v = 0" is equivalent to "minPTv divides
minPT". A moment's reflection gives:
Proposition. minPT is the least common multiple ofminPT v for all
Invoking Proposition 5.3.2 we obtain
Corollary. Every prime factor of%T is a factor o/minPT.
We shall see later (exercise ex5.3.8) that there are always vectors ν
such that minPT is equal to minPT v.
5.3.5 The minimal poynomial gives much information on Τ and on
polynomials in T.
Lemma. Let Px be a polynomial. Then Pl (T) is invertible if and only
ifPl is relatively prime to minPT.
PROOF: Denote Ρ = gcdf^, minPT). IfPY is relatively prime to minPT
then Ρ = 1. By Theorem A.6.2, there exist polynomials q, qx such that
qxPx +^minPT = 1. Substituting Τ for x we have q](T)Pl(T) = /, so
that Px (T) is invertible, and qx (T) is its inverse.
If Ρ φ 1 we write minPT = PQ, so that P(T)Q(T) = minPT(7) = 0
and hence ker(P(T)) D range(6(7)).
98
5. Invariant Subspaces
Since deg<2 < degminPT, the minimality of minPT guarantees that
Q(T) φ 0 so that range(^(7)) φ {0}, and since Ρ is a factor of Px,
ker(Px(T)) D ker(P(7)) D range(Q(T)) φ {0} and P{(T) is not in-
vertible. <4
Comments:
a. If Px (x) = x, the lemma says that Τ itself is invertible if and only if
ΓηίηΡτ(0)τ^0. The proof for this case reads: if m'\nPT=xQ(x), andTis
invertible, then Q(T) = 0, contradicting the minimality of minPT. On
the other hand, if minPT(0) = α φ 0, write R(x) = a~lx~l (a — minPT)
and observe that TR(T) = I — a~l minPT(T) = /, the identity, so that
R(T) = T~l.
b. If minPT is P(x), then minPT+;L, the minimal polynomial for Τ + Я,
is P(x — λ). It follows that Τ — λ is invertible unless χ — λ divides
minPT, that is, unless minPT(A) = 0.
5.3.6 The following proposition is important when F is not
algebraically closed and minPT = Φ is irreducible, but non-linear.
Proposition. Let Τ G «Sf(У) be such that Φ = minPT is irreducible
in F[x]. Then &>(T) = {P(T) : Ρ G ¥[x]} is afield.
PROOF: If Ρ e ¥[x] and P(T) φ 0, then gcd(P^) = 1 and hence P(T)
is invertible. Thus, every non-zero element in &(T) is invertible and
^(7) is a field. <
By Corollary 5.3.4, all the prime factors of %T are factors of, and
hence equal to, Φ = minPT. It follows that %T is a power of Φ. If
degΦ = d and %T = Фт then dim^ = deg#7 = md.
Ψ can now be regarded as an m-dimensional vector space Уф><Т\
over the extended field &(T) by treating the action of polynomials
P(T) on a vector ν as a multiplication of ν by the field element P(T). This
defines a system {Ψ^,^,Ί) in which a subspace of VmT\ is precisely
a 7-invariant subspace of Ψ.
The subspace spa η [7, ν], in У (over F) becomes "the line through ν
in (УыТ\)", i.e., the set of all multiples of ν by scalars from &(T)\ the
5.3. The minimal polynomial
99
statement "Every subspace of a finite-dimensional vector space (here
ψ over «^(Γ)), has a basis." translates here to: "Every ^-invariant
subspace of Ψ is a direct sum of cyclic subspaces, that is subspaces of
the form spa η [7, ν]."
EXERCISES FOR SECTION 5.3
ex5.3.1 Let Τ e &(T) and vGl Prove that if и е span[r, v], then minPTu
divides minPTv.
ex5.3.2 Let ^ be a ^-invariant subspace of У and Ty,^ the operator
induced on ψl<%. Let vGf, and let ν be its image in Ψj<ft. Prove that
minPT ~ divides minPTv.
ir/i^.v i,v
ex5.3.3 If (У, Γ) is cyclic (has a cyclic vector), then every S that commutes
with Γ is a polynomial in T. (In other words, &(T) is a maximal
commutative subalgebra of if (У).)
H/nf: If ν is cyclic, and Sv = />(Γ)ν for some polynomial P, then S = P(T).
ex5.3.4 a. Assume that minPTv = QXQ2, and let и — Qx(T)v. Prove that
minPT u - Q2.
b. Let Q e ¥[x] and assume that gcd(Q,minPT v) = Qv Write w = Q(T)v.
Prove that minPT w = Q2.
c. Assume that minPT v = PXP2, with gcd^,P2) = 1. Prove
span[7>] =span[7,,/>1(r)v]©span[r,/>2(r)v].
ex5.3.5 Let vx, v2 £ У and assume that minPT v and minPT v are relatively
prime. Prove that minPT v +v = minPT v minPT v .
Hint: Write P- = minPTv , β = minPTv +v , and let q- be polynomials such
that qxPx +q2P2 = 1. Then Qq2P2(T)(v{ + v2) = GOOO^) = 0, and so ^ Ι β.
Similarly P2 Ι β, hence /^ Ι β. Also, PlP2(T)(vl + v2) = 0, and βΙ/^.
ex5.3.6 Prove that every singular Τ £ а?(У) is a zero-divisor, i.e., there
exists a non-zero 5 G .Sf (У) such that ST = TS = 0.
H/nf; The constant term in minPT is zero.
ex5.3.7 Show that if minPT is divisible by Фт, with Φ irreducible, then there
exist vectors vGf such that minPT v = Фт.
ex5.3.8 Show that if a polynomial Ρ divides minPT, there exist vectors ν such
that minPT v = P. In particular, there exist vectors vGf such that minPT v =
minPT.
100
5. Invariant Subspaces
Hint: Use the prime-power factorization3 of minPT.
ex5.3.9 (У, T) is cyclic if and only if degminPT = dim^, i.e., if and only
if Τ is nonderogatory: minPT = χτ.
ex5.3.10 If minPT is irreducible, then minPT v = minPT for every ν φ 0 in Ύ.
ex5.3.11 If minPT is irreducible then dim У is divisible by degminPT.
Hint: Use Proposition 5.3.2.
ex5.3.12 Let PvP2e F[x]. Prove:
ker(P1(r))nker(P2(r)) = ker(gcd(/>1,P2)).
ex5.3.13 (Schur's lemma). A system {Ψ\У], У С У (Ψ), is minimal if
no nontrivial subspace of Ψ is invariant under every S e У.
Assume that {Ψ\У] is minimal, and Τ G ^(W).
a. If Τ commute with every S e У, so does P(T) for every polynomial P.
b. If Τ commutes with every 5еУ, then кег(Г) is either {0} or Ψ. That
means that Τ is either invertible or identically zero.
c. With Τ as above, the minimal polynomial minPT is irreducible.
d. If Τ commute with every ^У, and the underlying field is C, then
Τ = λΙ.
Hint: The minimal polynomial of Τ must be irreducible, hence linear.
ex5.3.14 Assume that Τ is invertible and degminPT = m. Prove that
minPT.! (x) = cxmm\r\PT(x~l),
where c = minPT(0)_1.
ex5.3.15 Let Τ e У{У). Prove that minPT vanishes at every zero of χτ.
Hint: If Tv = Av then minPTν=χ-λ.
Notice that if the underlying field is algebraically closed, the prime
factors of χτ are {χ — λ : λ £ σ(Τ)} and every one of these factors is minPT v,
where ν is the corresponding eigenvector. This is the most direct proof of
Proposition 5.3.2 (when F is algebraically closed).
ex5.3.16 Find the characteristic, and the minimal, polynomials of the 7 χ 7
matrix (α· ·) defined by
(l if 4 < 7 = 1+1 < 7,
ai i = λ
0 otherwise.
3SeeA.6.3.
5.3. The minimal polynomial
101
ex5.3.17 Assume that Λ is a non-singular matrix and let P(x) = Y^a^ be
its minimal polynomial. Prove that a0^0 зала show that Ρ gives an efficient
way to compute A~l.
Γθ -ll
ex5.3.18 Let A ,
R[x]} is isomorphic to C.
Hint: What is minPA?
e Л(1\Ж). Show that &>(A) = {P(A) : Ρ e
This page intentionally left blank
Chapter 6
Inner-Product Spaces
6.1 Inner products
Inner-product spaces, are real or complex vector spaces endowed
with an additional structure, called inner product. The inner product
allows for the introduction of geometric notions such as distance and
angle. Finite-dimensional real inner-product spaces are often called
Euclidean spaces. Complex inner-product spaces are also called
unitary spaces.
We shall have examples of both finite- and infinite-dimensional
inner-product spaces. However the theorems we prove assume, either
explicitly or implicitly, that the spaces involved are finite-dimensional.
6.1.1 Definition: An inner product on a real vector space Ψ is a
symmetric, positive definite bilinear form (·, ·) : Ψ χ Ψ —> R. In other
words, (·, ·) is a real-valued form satisfying
1. (w,v) is bilinear.
2. (w,v) = (v,w).
3. (и, и) > 0 for all иеУ, and (и, и) = 0 if and only if и = 0.
Examples:
a. The classical Euclidean η-space is W1 (also written £n), in which
(и, ν) = Σ^;> where и = [*j,... ,*„] and ν = [yx,... ,ул].
Z>. The space CE([0,1]) of all continuous real-valued functions on
[0,1], with the inner product: (/,g) = /0 f(x)g(x)dx.
103
104
6. Inner-Product Spaces
6.1.2 Definition: An inner product on a complex vector space ^
is a Hermitian, positive definite, sesquilinear form (·,·): У x ^ —> C.
In other words, (·, ·) is a complex-valued form satisfying
1. (и, ν) is sesquilinear, that is, linear in и and skew linear in v:
(Am,v) = Я(м,у) and (m,Av) = λ (и, ν).
2. (и, ν) is Hermitian, that is, (и, v) = (v,m).
3. (и, и) > 0, with (и, и) = 0 if and only if и = 0.
Notice that the sesquilinearity follows from the Hermitian
symmetry combined with the assumption of linearity in the first entry.
Examples:
a. In Cn we define the inner product of the vectors и = [χχ,... ,xn]
and ν = [уj,... ,yn] by (и, ν) = Е*,\У/· If we consider the vectors
IV
as columns, и =
multiplication).
Xn.
and ν =
Un.
, then (m,v) = ν'и (matrix
b. The space C([0,1]) of all continuous complex-valued functions on
[0,1]. The inner product is defined by (f,g) = /0 f(x)g(x)dx.
We shall reserve the notation Ж for inner-product vector spaces,
whether real or complex, and to avoid duplication we write the
formulae corresponding to complex Ж, remembering, however, that
complex conjugation is well defined on Ш (as the identity).
The distinction between real and complex Ж will be explicit when
the results depend on the underlying field being algebraically closed,
which С is, and К is not.
6.1.3 Given an inner-product space Jff, we define a norm on it by
(6.1.1) M = vw>-
Lemma (The Cauchy-Schwarz inequality).
(6.1.2) |<a,v)|<NIN|.
6.1. Inner products
105
PROOF: If ν is a scalar multiple of u, we have equality. If v, и are not
proportional, then, for IgR,
0< (w + Av,w + Av) = ||w||2 + 2A9t<>,v) + A2||v||2.
The right-hand side is a quadratic polynomial (in the variable A) with
real coefficients and no real roots, which means that its discriminant,
(91(m,v))2-||w||2||v||2, is negative, that is, |91(m,v>| < ||u||||v||. This
completes the proof for real Ж.
If Ж is complex: we have |9l(£w,v)| < |H|||v|| for every ζ G С
with \ζ\ = 1. Take ζ such that (£m,v) = |(m,v)|. <
6.1.4 The norm has the following properties:
a. Positivity: ||0|| =0; if ν φ 0 then || v|| > 0.
b. Homogeneity: ||αν|| = |<z|||v|| for scalars a and vectors v.
c. The triangle inequality: ||v + m|| < ||v|| + ||m||.
d. The parallelogram law: ||v + m||2 + ||v-m||2 = 2(||v||2 + \\u\\2).
Properties a and b are obvious. Property с is equivalent to
||v||2+||w||2 + 29l(v,w)< ||v||2 + ||m||2 + 2||v||||ii||,
which reduces to (6.1.2). The parallelogram law is obtained by
"opening brackets" in the inner products that correspond to ||v + m||2 and
l|v-«||2·
The first three properties are common to all norms, whether
defined by an inner product or not. They imply that the norm can be
viewed as length, and δ (и, ν) = \\u — v\\ has the properties of a metric.
The parallelogram law, on the other hand, is specific to, and in fact
characteristic of, the norms defined by an inner product.
Proposition. A norm defined by an inner product determines the
inner product.
The proof is left to the reader as exercises ex6.1.14 and ex6.1.15.
I об 6. Inner-Product Spaces
6.1.5 Orthogonality. Let Ж be an inner-product space.
Definition: The vectors v.uinJif are said to be (mutually)
orthogonal, denoted ν J_ u, if (v,u) = 0. Observe that, since (m,v) = (v,m),
the relation is symmetric: и J_ ν <=> ν J_ w.
The vector ν is orthogonal to a set А с Ж, denoted ν J_ A, if it is
orthogonal to every vector in A.
If both ν J_ A and и J_ A then, for every w eA,
(αν + few, w) = a(v, w) + b(u, w) = 0.
It follows that for any set А с Ж, the set AL = {ν: ν J_ A} is a subspace
of Ж. (This notation is consistent with 3.1.3; see 6.2.1 below).
Similarly, if ν J_ A, and both wl and w2 are in A, then we have
(v,aw] +bw2) = <z(v,Wj) +fe(v,w2) = 0
so that ν J_ (spaη [A]). In other words: A1- = (spaη [A])-1.
A vector ν is said to be normalif ||v|| = 1. A sequence {v],..., vw}
is orthonormal if
(6.1.3) (ν,.,νj) = S, . (i.e., 1 if i = ;, and 0 if *y ;);
that is, if the vectors ν · are normal and pairwise mutually orthogonal.
Propositions. Let {u],..., um} be orthonormal in an n-dimensional
space Ж.
a. {u],..., um} is linearly independent.
PROOF: If Y^cijUj — 0 then ak = (£я -и -,ил) = 0 f°r a^ * ^ [1> ml· ^
b. For every ν G ^f, ί/ге vector v{ = ν — Σ™ (ν, и -)и · /s orthogonal to
span[uv...,um].
PROOF: (v]5^) = (ν,ίΐ^) — (ν,ιΐβ) = 0 for every к G [1, m]. If w is a
vector in spanf^,...,ит], w = Yl\Ckuk, and it follows that
<νΙ}νν> = Σ7·^<ν,,«Λ>=0. *
с If{u^...^un}isan orthonormal basis and ν G Жу then
η
(6.1.4) v = Yi{v,Uj)uj.
1
6.1. Inner products
107
PROOF: If spanfwj,... ,ып] = Ж, then ν — L" (v,m -)и · is orthogonal to
itself, and so ν = Σ" (ν, иу-)иу·. ч
d. ParsevaVs identity. If{uv..., un] is an orthonormal basis for Жу
then for all v, w G ^,
«
(6.1.5) (ν, w) = £(v, и,·) (w, w;·) .
ι
proof:
= Ey<v,My.>(w,Mi/> ч
e. BesseVs inequality and identity. If {ux,...,um} is orthonormal
and ν G c^f, i/zen
(6.1.6) Il(v,«;-)|2<||v||2.
If{ux,.. . ,и„} is an orthonormal basis for Ж> then ||v||2 = Y!\\(v^uj)\2-
PROOF: The identity, when {ир...,ит} is a basis, is clearly a
special case of (6.1.5). The inequality follows from this and the fact
(see Corollary 6.1.6 below) that every orthonormal {uv...,um} can
be completed to an orthonormal basis. <
f. The matrix {ai ) = AT v of an operator Τ G J£(Jf) relative to an
orthonormal basis ν = {vx,..., vn}, is given by at = (Tv ·, v·).
PROOF: The fth entry in the /th column of AT v is the coeffcient a{
of vt in the expansion of Tv- in the basis v; see subsection 2.4.3. If ν is
orthonormal and Τ G Jf(Jff), then, by (6.1.4), wehave<2/; = (7V·,vt).
6.1.6 Proposition. (Gram-Schmidt). Let {v1,...,vw} be
independent; then there exists an orthonormal {iip... ,nm} such that for all
ke [l,m],
(6.1.7) span [и,,... ,ик] = span[vp..., vj.
108
6. Inner-Product Spaces
Proof: (By induction on m). The independence of {vp..., vw}
implies that Vj Φ 0. Write ux = v1 /Hvj||. Then ux is normal and (6.1.7)
is satisfied for к = 1.
Assume that {ux,..., щ} is orthonormal and that (6.1.7) is satisfied
for к < I. Since v/+1 0 spanfjvj,..., vz}], the vector
v/+i=v/+1-£(v/+1,My>y.
y=i
is nonzero and (by Lemma 6.1.5, part b) is orthogonal to u- for all
j<l. WesetM/+1=v/+1/||v/+1||. <
Corollary. Every orthonormal sequence {ux,...,um} in Ж can be
completed to an orthonormal basis. Hence, every finite-dimensional
Ж has an orthonormal basis.
Proof: Observe that {u],..., um} is independent, complete it to a
basis, apply the Gram-Schmidt process and notice that it does not change
the vectors Uj,l<j<k. <
6.1.7 If Ψ с Ж is an m-dimensional subspace, and {v ■}? is a basis
for Ж such that {v^™ is a basis for Ψ, then the basis {и,·}" obtained
by the Gram-Schmidt process splits into two: {uj}™ U {и ·}^+1, where
{иЛ™ is an orthonormal basis for W and {uJ^+i is an orthonormal
basis for WL. This gives a direct sum (in fact, orthogonal)
decomposition Ж = Ψ θ WL.
The map
m
(6.1.8) πψ\ v^J^(v,Uj)uj
ι
is called the orthogonal projection onto W.
The orthogonal projection onto W depends only on W and not on
the particular basis we started from. In fact, if ν = vx + v2 = u] + u2
with Vj and ux in W, and both v2 and u2 in WL, we have
v, -U]=u2-v2 Gfnf1,
which means that vl—u] = u2 — v2 = 0, that is, vx = Mj and w2 = v2.
6.1. Inner products
109
6.1.8 The definition of the distance δ(ν1, v2) (= \\v] — v2||) between
two vectors, extends to that of the distance between a point (v G Ж)
and a set (E с Ж) by setting δ(ν,Ε) = mfueES(v,u).
The distance between two sets, E] and E2 in Ж, is defined by
(6.1.9) δ(ΕνΕ2) = Щ\\у{ -v2|| : VjtEj}.
The proof of the following proposition is left as exercise ex6.1.5.
Proposition. Let W С Ж be α subspace, and ν G Ж. Then
δ(ν,Ψ) = \\ν-πΨν\\.
In other words, π^ν is the vector closest to ν in W.
EXERCISES FOR SECTION 6.1
ex6.1.1 Let Ψ be a finite-dimensional real or complex vector space, and
ν = {νλ,..., vn) a basis. Explain: "declaring {vj,..., v„} to be orthonormal
defines an inner product on У'.
Hint: Define (·, -)v by: for и = Σα/ν/ and w = Σ^/ν/> set (и, w)v = Σ^Α-
ехб.1.2 Prove that if Ж is a complex inner-product space and Τ G j£f (^),
there exists an orthonormal basis for Ж such that the matrix of Τ with respect
to this basis is triangular.
Hint: See Corollary 5.2.8.
ex6.1.3
a. Let Ж be a real inner-product space. The vectors v, и are mutually
orthogonal if and only if ||v +w||2 = ||v||2 + ||w||2.
b. If Ж is a complex inner-product space, and v, и G Ж, then the condition
||v + w||2 = ||v||2 + ||w||2 is necessary, but not sufficient, for ν JL u.
Hint: Consider the case "(w, ν) purely imaginary".
c. If Ж is a complex inner-product space, and v, и G Ж, the condition: For
all a, beC, \\av-\-bu\\2 = |a|2||v||2 + |£|2||w||2 is necessary and sufficient for
ν JL u.
d. Let У and <% be subspaces of Ж. Prove that У JL ^ if and only if for
allvG^ andwG ^, ||v + w||2 = ||v||2 + ||w||2.
e. The set {v1?...,vw} is orthonormal if and only if ||Εα·ν·||2 = Σΐα;·|2 f°r
all choices of scalars a-J= 1,..., m. (Here Ж is either real or complex.)
по
6. Inner-Product Spaces
ехб.1.4 Show that the map πψ defined in (6.1.8) is an idempotent linear
operator (that is, ττ^ = πψ) and is independent of the particular basis used
in its definition.
ex6.1.5 Prove Proposition 6.1.8.
ex6.1.6 Let Ε- = ν · + Ψ- be affine subspaces in Ж. What is δ(Εχ,Ε2)Ί
ехб.1.7 Show that the sequence {ux,..., um} obtained by the Gram-Schmidt
procedure is essentially unique: each u- is unique up to multiplication by a
number of modulus 1.
Hint: If {vj,...,vw} is independent, and Wk = span[{v1?... ,v^}], 0<k <m,
then uj is cnw±vj, with \c\ = \\πψΣν]\\~χ.
ехб.1.8 Let Α, Β £ <M(n, C). Prove, by reference to the standard inner
product on Cn , that (А,Б) = trace(A£ ') is an inner product on Л(п, С).
ехб.1.9 Let A^ Ж(п\ С) and assume that its rows w ·, considered as vectors
in Cn, are pairwise orthogonal. Prove that AA r is a diagonal matrix, and
conclude that |detA| = П1ку-||·
ехб.1.10 Let {v{,..., vn} С Сп be the rows of the matrix A. Prove Hadamard's
inequality:
(6.1.10) |detA|<ni|vy-||.
Hint: Write Wk = span[{vl5... ,у^}]Д = 0,...,«- 1, w·■, = Kw± v-, and
apply the previous problem.
ex6.1.11 Let v,v'\u,u' e Ж. Prove
(6.1.11) |(У)И)_(/)И')|<||У||||И_И'|| + ||И'||||У_У'||.
(Observe that this means that the inner product (v, u) is a continuous function
of ν and и in the metric defined by the norm.)
ex6.1.12 The standard operator norm1 on ££(Ж) is defined by
(6.1.12) ||Г|| = тах||Гу||.
N1=1
Let A £ JK(n\C) be the matrix corresponding to Τ with respect to some
orthonormal basis and denote its columns by и ·. Prove that
(6.1.13) sup||«;||<||rH<(f||M;||2)^.
1See*2.6.
6.2. Duality and the adjoint
111
ex6.1.13 Let Ψ С Ж be a subspace and π a projection on Ψ along a sub-
space Ж. Prove that ||π|| = 1 if and only if Ψ = Ух, that is, if π is the
orthogonal projection on Ύ.
ex6.1.14 Prove that in a real inner-product space, the inner product is
determined by the norm: (polarization formula over R)
(6.1.14) (M)V) = I(||K + V||2_||1<_v||2)i
ex6.1.15 Prove: In a complex inner-product space, the inner product is
determined by the norm; in fact (polarization formula overC)
(6.1.15) (w,v) = i(||w + v||2-||w-v||2 + /||w + /v||2-/||w-/v||2).
ex6.1.16 Show that the polarization formula (6.1.15) does not depend on
positivity: given a sesquilinear Hermitian form ψ (on a vector space over C),
define the Quadratic form associated with it by: Q(y) = i//(v, v). Prove
(6.1.16) y/(w, v) = - (Q(u + v) - Q(u - v) + iQ(u + iv) - iQ(u - iv)).
ex6.1.17 Let R,T e <£{Ж). If (Rv, ν) = (Γν, ν) for all vE-Ж, then R = T.
Hint: Check that (flv,u) = (Γν,и) for all у.иеЖ.
6.2 Duality and the adjoint
6.2.1 Ж as its own dual. The inner product defined in Ж
associates with every vector и G Ж the linear functional φ14: ν н-> (ν, и). In
fact every linear functional is obtained this way.
Theorem. If φ is α linear functional on a finite-dimensional inner-
product space Ж, then there exists a unique w G Ж such that for all
уеЖ,
(6.2.1) 9(v) = ^(v) = (v,w>.
PROOF: Let {иЛ be an orthonormal basis, and let w = Σ ф(м/)м/· F°r
all 7, we have <р(и ·) = (w, w ·). For ν G ^f, ν = Σ(ν, w )w , hence
112
6. Inner-Product Spaces
φ(ν) = Σ(ν,^·)φ(^·) = I(v,My.)(w,Mi/>
and by Parseval's identity this equals (v,w). ^
In particular, an orthonormal basis in Ж is its own dual basis.
6.2.2 The adjoint of an operator. Once we identify Ж with its
dual space, the adjoint of an operator Τ G Л£{Ж) is again an operator
on Ж. Indeed,2 given и G Ж, the mapping ν н-> (TV, и) is a linear
functional and therefore equal tov^ (v, w) for some w G ^f. We
write T*u = w and check that ыу-^w is linear. In other words, 7* is a
linear operator on ^f, characterized by
(6.2.2) (Γν,Μ> = (ν,Γ*Μ>.
Lemma. For Τ G JzfpT), (Γ*)* = T.
PROOF: (v, (Г*)*и) = (Γ*ν,и) = (и,Г*ν) = (Γιι,ν) = (ν,7м). Ч
Proposition 3.2.4 reads in the present context as
Proposition. For Τ G JSf (Ж\ range(r) = (кег(Г*))±.
PROOF: Since (Tx,y) = (x,T*y), we have у J_ range(r) if, and only
if T*y J_ Ж, that is if and only if ye кег(Г*). ч
6.2.3 The adjoint of a matrix.
Definition: The adjoint of a matrix A e Ж{п\ С) is the matrix A* =
A*. The matrix A is self-adjoint, or Hermitian, if A = A*, i.e., if α·; = я^
for all ij.
Notice that for matrices with real entries the complex
conjugation is the identity, the adjoint is the transposed matrix, and self-adjoint
means symmetric.
Let (α· ·) = AT v be the matrix of an operator Τ relative to an
orthonormal basis v, and (£· ·) = A7* v the matrix of 7* relative to the
same basis. By Proposition f of 6.1.5, we have a{ · = (Tv ·, v·) and
2This repeats the argument of section 3.2 in the current context.
6.3. Self-adjoint operators
113
bU = (T*VP Vi> = (VP ТУг) = 4ϊ lt f°ll0WS that A7V = (ΑΓ,ν)* ОГ>
in words: The matrix of the adjoint is the adjoint of the matrix.
In particular, Τ is self-adjoint if and only if AT v is self-adjoint, for
every orthonormal basis v.
EXERCISES FOR SECTION 6.2
ex6.2.1 Prove that if T,S G а?(Ж), then (ST)* = T*S*.
ex6.2.2 Prove that if Τ G <£(Ж\ then ker(7*7) = кег(Г).
ехб.2.3 Prove that the characteristic polynomial χτ* is the complex
conjugate of χτ.
ехб.2.4 Show that if Tv = Av, T*u = μ«, and μ φ λ, then (ν, и) = 0.
ехб.2.5 Show that if Я = а + bi, α, b G IR, then
||(Γ-λ)ν||2 = ||(Γ-α)ν||2 + |^|2||ν||2.
ехб.2.6 Rewrite the proof of Theorem 6.2.1 along the following lines: If
ker(<p) = Ж, then φ = 0 and u* = 0. If not, dimker(<p) = dim Ж - 1 and
(ker((jp))± ^ 0. Take any nonzero й G (кег(<р))х and set и* = см where the
constant с is the one that guarantees (й, ей) = φ(ϋ), that is, с = ||й||-2<р(й).
6.3 Self-adjoint operators
6.3.1 Recall that an operator Τ G Л£(Ж) is self-adjoint if 7*
coincides with 7, that is, if (7w, ν) = (и, Г ν) for every m,vG ^.
Examples:
a. For Τ G ££(Ж) (real or complex), the operators 77* and 7*7
are both self-adjoint.
b. For an arbitrary operator 7 G Л£(Ж), ^f complex, the operators
(6.3.1) 917 = ^(7 + 7*) and 37 = ^(7-7*)
are both self adjoint (check!). They are called the real and imaginary
parts of 7, respectively, and we note that 7 = 917 + /37. If Ж is real,
then 917 is self-adjoint, but 37 is not defined.
с If 5 and 7 are self-adjoint then so is 5+ 7. In particular, if 7 is
self-adjoint and λ G Μ then 7 — λ is self-adjoint.
114
6. Inner-Product Spaces
The structure of self-adjoint operators follows from the following
observations.
Proposition. Assume that Τ is self-adjoint on Ж. Then
a. σ(Τ) С R.
b. If Ψ а Ж is Τ -invariant, then so is Ψ1-.
c. If"W С Ж is Τ-invariant, then the restriction Τ ψ of Τ to W is
self-adjoint.
Proof:
a. If λ G o{T) and ν is a corresponding eigenvector, then
Я ||v||2 = <rv, v> = <v, rv> = Я ||v||2, so that λ = 1.
b. If ν G W1-, then, for every w G W, we have Tw G W and hence
(Γν, w) = (v, Tw) = 0. It follows that Tv G WL.
c. The condition {Twx, w2) = (w,, 7w2) is valid when w G Ж, since
it holds for all vectors in Ж. .
6.3.2 The previous proposition shows that a self-adjoint operator Τ
induces an orthogonal decomposition of Ж into 7-invariant subspaces.
The spectral theorem shows that this decomposition can be refined so
that all of the subspaces be 1-dimensional.
Theorem (The spectral theorem for self-adjoint operators). Let
Ж be a finite-dimensional, inner-product space, and let Τ G ££(Ж)
be self-adjoint. Then there is an orthonormal basis {u],..., un} of Ж,
each element of which is an eigenvector of Τ\
PROOF: The proof is by induction on η = dim Ж. lfn= 1, then every
vector in Ж is an eigenvector of 7, and the statement is obvious.
Assume that η > 1 and that the statement is true for m<n. Because
o(T) С М, it follows that <з(Т) is nonempty for both real and complex
6.3. Self-adjoint operators
115
inner-product spaces. Thus, Τ has an eigenvector un G Ж, and we
may assume that ||ил|| = 1.
Since span[w„] is Γ-invariant, it follows that Ж1 = spanf^]-1 is
also T-invariant, by part b of the proposition above. By part с, 7\^,
is self-adjoint, so by the induction hypothesis, there is an orthonor-
mal basis {u],..., un_]} of Ж1 consisting of eigenvectors of T. Then
{ux,..., un] is the required orthnormal basis of Ж. <
Note that the matrix of Τ relative to a basis of eigenvectors is diagonal.
6.3.3 If Τ is an operator on an inner-product space Ж and A G σ(Γ),
then we denote by Жх the kernel of Τ — A. That is,
(6.3.2) Жх = кег(Г - A) = {ν G Ж : Tv = Αν}.
Жх is called the eigenspace of Τ corresponding to A. We denote by
πλ the orthogonal projection on Жх.
With this notation we can restate the spectral theorem, Theorem
6.3.2, in the following form. Some authors refer to this as the spectral
theorem for self-adjoint operators.
Theorem. Let Ж be an inner-product space and Τ a self-adjoint
operator on Ж. Then Ж = ®хео(т)^х> where Ж^ J_ Ж^ when
Aj φ λ^, and
(6.3.3) τ= Σ λπλ·
λ£σ{Τ)
The decomposition Ж = (Βχ£σ(Τ\ ^χ is often referred to as the
spectral decomposition induced by Τ on Ж. The representation (6.3.3)
is called the spectral decomposition of T.
6.3.4 If {u],..., un] is an orthonormal basis whose elements are
eigenvectors for 7\ say Tuj = Aw, then
(6.3.4) Tv = ^j(v9uj)uj
for all ν G Ж. Consequently, writing a- = (v,m ·) and ν = Σα.и-,
(6.3.5) <7V,v) = 2».|2 and ||7v||2 = £|A/|(v,M;)|2.
116
6. Inner-Product Spaces
Proposition. Assume that Τ is self-adjoint, then || Τ || = тахЯе ,^ | A |.
PROOF: If Xm is an eigenvalue with maximal absolute value in σ(Γ),
then || Τ || > \\Tum\\ = max^^JAI. Conversely, by (6.3.5),
\\Tv\\2 = ЦЯ/|(у,^.)|2 < max|A/£|(v^.)|2 = max|A/||v||2.
6.3.5 Commuting self-adjoint operators. Let Τ be self-adjoint,
and let Ж = (ΒχΕσ(Τ\ ^χ be the spectral decomposition induced by
T. If 5 commutes with T, then 5 maps each Жх into itself. Since
the subspaces Жх are mutually orthogonal, if 5 is self-adjoint then
so is its restriction to each Жх, and we can apply Theorem 6.3.3 to
every one of these restrictions and obtain, in each Ж^, an orthonormal
basis made up of eigenvectors of 5. Since every vector in Жх is an
eigenvector for 7\ we obtained an orthonormal basis, each of whose
elements is an eigenvector both for Τ and for 5. We now have the
decomposition
λΕσ(Τ),μΕσ{Ξ)
where Жх = ker(7 - Α) Π ker(5 - μ).
If we denote by πχ the orthogonal projection onto Жх , then
(6.3.6) τ = Σλπλ,μ and s = Im%-
More generally, given a set of commuting, self-adjoint operators
{7}}7=i С &{Ж) and eigenvalues λ- e σ(7}), we set
m
^„..„я^ПМ^-я,).
7=1
The joint spectrum of {ТЛу=1 is the set of m-tuples
а(Г„...,7'и) = {(Я1,...,Ада):^1___)Дт^{0}}.
6.3. Self-adjoint operators
117
If we denote by πχ , the projection onto Жх , , then, as above,
we have Ж = , and
Aj ,...,ЛШ
(6.3.7) Tj= Σ ЯА,..Ат, forl<;<m.
(λ,,...Λη)Εσ(η,...,7;η)
The spaces Жх λ are clearly invariant under every operator
in the algebra generated by {Τ-}™=], that is, operators of the form
P(TX,..., 7W), where Ρ is a polynomial in m variables. These
observations yield the following generalization of Theorem 6.3.3.
Theorem. Let Ж be a finite-dimensional inner-product space, and
let {T-}™={ be commuting self-adjoint operators on Ж. Let srf С
«if (Ж) be the subalgebra generated by {Tj}™=v Then
(6.3.8) Ж= 0 ЖХху^
(Я,,...Лп)еа(7Р...,7;п)
where Ж^^ -L J^p...,Mm ϊ/(Α,,... ,AW) φ (др...,дт), and every
S ^ si is a scalar operator on each <3ή?λ λ . To be more precise, if
S = P(TV... ,Гт), then
(6.3.9) 5= £ ^ρ-ΛΚ ^·
(Я1,...,ДЯ1)еа(Г1,...,Гт)
The verification of equation (6.3.9) is left as exercise ex6.3.3.
Every vector in Жх , isa common eigenvector of all the
operators 5 G si. If we choose an orthonormal basis in every Μ'χ λ , the
union of these is an orthonormal basis of Ж with respect to which the
matrices of all the operators in stf are diagonal.
6.3.6 If Ж is complex, then a subalgebra si С ^£{Ж) is self-adjoint
ifSes/ implies that 5* G si. By considering real and imaginary parts,
we see that a self-adjoint algebra is generated (in fact spanned) by the
self-adjoint elements it contains.
With this terminology, Theorem 6.3.5 may be restated as follows.
118
6. Inner-Product Spaces
Theorem. If Ж is a complex inner-product space, and szf is a
commutative, self-adjoint subalgebra of ^£(Ж\ then there is an orthonor-
mal basis {u],... ,un} of Ж such that every uk is an eigenvector of
every Тез/.
EXERCISES FOR SECTION 6.3
ex6.3.1 Τ e а?(Ж) is self-adjoint if and only if the quadratic form
Q(v) = (TV, v) is real-valued.
ex6.3.2 Let T,S G ££(Ж) be commuting, self-adjoint operators. Show that
P(S, T) is a self-adjoint operator for every polynomial Ρ with real
coefficients.
ехб.3.3 Verify equation (6.3.9).
ex6.3.4 Let Τ e *5?(Ж) be self-adjoint, let λλ < λ^ < · · · < λη be its
eigenvalues and {иЛ the corresponding orthonormal eigenvectors. Prove the "min-
max principle":
(6.3.10) A, = min max (Tv.v).
' dim^=/ vg^.||v|| = 1
H/nf: Every /-dimensional subspace intersects span [{μ ,·}"=/] see 1.3.7.
ехб.3.5 Let W С Ж be a subspace, and тг^ the orthogonal projection onto
Ψ. Prove that if Τ is self-adjoint on Ж, then πψΤ is self-adjoint on Ж.
ехб.3.6 Let Л £ Ж(п^Ж) be symmetric. Prove that #A has only real roots.
ex6.3.7 Assume that A £ <M{m, С) is Hermitian. Let ζ denote the column
vector with complex entries z1?...,zm, and assume that (Λζ,ζ) > 0 for all
choices of 7 . Prove that all the eigenvalues of Λ are nonnegative.
Show that if (Az, z) > 0 for all ζ φ 0, then detA > 0.
ex6.3.8 Find an operator Τ £ j£f (M2) that commutes with its adjoint, but has
no eigenvectors. How does this example relate to Theorem 6.3.6?
Hint: Consider rotations.
ex6.3.9 Show that an algebra over С is self-adjoint if and only if it is
generated by self-adjoint operators.
ex6.3.10 Let SS be a commutative self-adjoint subalgebra of ££(Ж). Prove
that:
a. The dimension of 38 (over C) is bounded by а\тЖ.
6.4. Normal operators
119
b. 38 is generated by a single self-adjoint operator, i.e., there is an operator
IGJ such that 3S = {P(T) : Ρ G C[x]}.
c. Зё is contained in a commutative self-adjoint subalgebra of j£f (Ж) of
dimension а\тЖ.
ехб.3.11 The Gram determinant detT of the vectors ν , j = 1,..., m, in Jf7
is the detrminant of the matrix
r = r(v1?...,vw)
(vpVj) (v1?v2) (v1?vw)'
|_(ут^) (vw,v2) (vm,vm)J
a. Identify the inequality r(vj, v2) > 0.
b. Prove that detT > 0, and that it vanishes if and only if the vectors ν are
linearly dependent.
Hint: Let ζ denote, as above, the column vector with entries ζλ, ·.. ,zm- Show
that
m
(6.3.11) (Γχ,χ)Η|Σ>,ν,||2.
6.4 Normal operators
We assume in this section that Ж is a complex inner-product
space.
6.4.1 Definition: An operator 7 e ^f(Jif) is norma/ if it
commutes with its adjoint, i.e., if 77* = 7*7.
Self-adjoint operators are clearly normal. If 7 e Л?(Л?) is normal,
then 5 = 77* = 7*7 is self-adjoint.
We observe that 7 is normal if and only if its real and imaginary
parts, 917 and 37, commute (see example b of 6.3.1). Furthermore,
exercise ex6.4.3 below implies that if 5 and 7 are both normal and
commute, then 35, 915, 37 and 917 all commute.
These observations and exercise ex6.3.9 together imply the
following lemma.
120
6. Inner-Product Spaces
Lemma. If {Τλ,..., Tm} are commuting normal operators, then the
algebra that they generate is contained in a commutative, self-adjoint
algebra.
Theorem 6.3.6 therefore implies the spectral theorem for normal
operators.
Theorem (The spectral theorem for normal operators). Let Ж be
a finite-dimensional, complex inner-product space and srf С Л£(Ж)
a commutative subalgebra generated by normal operators. Then there
exists an orthonormal basis {uk} of Ж such that every uk is an
eigenvector for every ГЕ^.
In particular, IfT£ ^(Ж) is normal, then Ж has an
orthonormal basis of eigenvectors ofT.
EXERCISES FOR SECTION 6.4
ex6.4.1 Prove without using the spectral theorems:
a. If S is normal, then ker(S) = ker(S*).
b. If S is normal, then ker(S) = ker(S2).
c. If S is normal and Sv = Av, then S*v = A v.
d. S is normal if and only if ||5*v|| = ||5v|| for all vGJf.
Hint: See exercise ex6.1.17.
ex6.4.2 Prove: if S is normal, then a necessary and sufficient condition for an
operator R to commute with S is that all the eigenspaces of S be 7?-invariant.
ex6.4.3 Show that if S is normal and R commutes with S, then R commutes
with S* as well.
ex6.4.4 If Τ is normal and σ(Τ) С Μ, then Τ is self-adjoint.
ехб.4.5 Show that if S is normal, then S and S* have the same
eigenvectors with the corresponding eigenvalues complex conjugate. In particular,
o(S*)=^{S).
ex6.4.6 Let Б be a commutative self-adjoint subalgebra of ££(Ж). Prove:
a. The dimension of Б is bounded by а[тЖ.
b. В is generated by a single self-adjoint operator, i.e., there is an operator
Τ G В such that В = {P(T) : Ρ G C[jc]}.
6.5. Unitary and orthogonal operators
121
с. В is contained in a commutative self-adjoint subalgebra of ££(Ж) of
dimension а\тЖ.
6.5 Unitary and orthogonal operators
We mentioned in subsection 6.1.4 that the norm in Ж defines a
metric, the natural metric for which the distance between the vectors ν
and и is given by δ(ν,ιι) = ||v —и||.
Maps that preserve a metric are called isometries of the given
metric. A particular class of isometries for the natural metric on Ж are the
linear isometries, that is, operators U G Л£(Ж) such that \\Uv\\ = ||v||
for all ν G Ж. These are called unitary operators when Ж is complex,
and orthogonal operators when Ж is real. The operator U is unitary if
\\Uv\\2 = (t/v,t/v> = (v,C/*C/v) = (v,v>.
It is an easy exercise to verify that polarization (see ex6.1.15) extends
the equality (ν,ί/*ί/ν) = (v,v) to
(6.5.1) (m,C/*C/v) = (m,v),
for all и, ν G У, which implies that U*U = /. Since Ж is assumed
finite-dimensional, a left inverse is an inverse, so [/* = [/-!. Observe
that this implies that unitary operators are normal, and the spectral
theorem for normal operators is valid for algebras of commuting unitary
operators.
6.5.1 Proposition. Let Ж be an inner-product space, Τ G ££{Ж\
The following statements are equivalent:
a. T is unitary;
b. Τ maps some orthonormal basis onto an orthonormal basis;
c. Т maps every orthonormal basis onto an orthonormal basis.
PROOF: If Τ is unitary, then it maps orthonormal sequences to
orthonormal sequences, which implies parts b and с
Assume that both {ux,..., un} and {Tul,..., Tun} are orthonormal
bases. For any ν G Ж, we have ν = Σβ,Μ, an<3 Tv = Σα Tu:, and by
Bessel's identity, ||7"v||2 = Σ|β/|2 — llvl|2> so Τ is unitary. <
122
6. Inner-Product Spaces
The columns of the matrix of a unitary operator U relative to an or-
thonormal basis {v ·} are the coefficient vectors of Uv ■ and, by (6.1.5)
(Parseval's identity), are orthonormal in Cn (respectively W1). Such
matrices (with orthonormal columns) are called unitary when the
underlying field is C, and orthogonal when the field is R.
The set <9ί(η) С JK{n\C) of unitary nxn matrices is a group
under matrix multiplication. It is caled the unitary group.
The set G(n) С Ж(п\Ж) of orthogonal nxn matrices is a group
under matrix multiplication. It is caled the orthogonal group.
6.5.2 Unitary equivalence. Recall that А,В G JK{n\¥) are similar,2,
i.e., they represent the same operator relative to (possibly) different
bases, if and only if В = С-1 AC for some С G GL(n,F). In other
words, two matrices are similar if they are conjugate under the
action of GL(n,F), where the conjugating matrix maps one basis onto
another.
In the case that the vector space is a real or complex inner-product
space, and if we want the conjugation to preserve geometric
information, the two bases have to have the same geometry: corresponding
basis elements must have the same length, and inner products between
pairs of corresponding basis elements must be the same. This means
that we have to restrict the choice of conjugating matrices to those that
preserve the geometry, that is, to unitary matrices for a complex inner-
product space, and to orthogonal matrices when the inner-product
vector space is real.
Definition: The matrices А, В G JK{n\ C) are unitarily equivalent if
there exists a matrix U G ^/(n) such that A = U~]BU.
The matrices А, В G Ж(п\Ж) are orthogonally equivalent if there
exists a matrix О G 6{n) such that A = 0~lBO.
Consider for example a matrix A G JZ(n\C) that has η distinct
eigenvalues. Then A has η independent eigenvectors which can be
taken as a basis, and changing the standard basis to that basis
conjugates A to a diagonal matrix. The matrix A is similar to a diagonal
3See Proposition 2.4.6.
6.5. Unitary and orthogonal operators
123
matrix, but is unitarily equivalent to one only if its eigenvectors are
pairwise orthogonal.
On the other hand, applying the Gram-Schmidt procedure to the
basis obtained in the course of the proof of Corollary 5.2.8, one obtains
the proof of the following theorem.
Theorem. Every matrix A £ ^(n;C) is unitarily equivalent to an
upper triangular matrix.
A representation A = U~]BU, where U is unitary and В upper
triangular, is called a Schur decomposition of A. We note that the matrices
В and U are not unique (see exercise ex6.5.6).
6.5.3 Spectral theorem for Hermitian/symmetric matrices. Let
A £ «y#(n;C), and let TA denote the operator of left multiplication by
A on C". The operator TA is self-adjoint if and only if A is Hermitian.
In this case, Theorem 6.3.3 guarantees that C" has an orthonormal
basis {v } all of whose elements are eigenvectors for the operator of
left multiplication by A.
The matrix of TA relative to the basis {v,,..., vn} is diagonal, and
the matrix С that affects the change of basis maps the orthonormal
basis {vj,..., vn} onto the standard basis of C" and hence is unitary.
This proves the following theorem:
Theorem. Every Hermitian matrix in Ж{п\ С) is unitarily equivalent
to a diagonal matrix.
6.5.4 If A £ Ж(п\Ж) is symmetric, then the same reasoning applies
to prove the following theorem:
Theorem. Every symmetric matrix in Ж{п, Ж) is orthogonally
equivalent to a diagonal matrix.
6.5.5 Harmonic Analysis for finite abelian groups. Consider a
finite abelian4 group G, of order \G\. We write the group operation as
4The generalization of the following ideas to nonabelian groups is called
representation theory. We give a brief survey of some of the basic elements of the
representation theory for finite groups in Section 8.5.
124
6. Inner-Product Spaces
addition. Denote by C(G) the \G\-dimensional complex vector space
of all complex-valued functions on G. Define an inner product in C(G)
by: for φ, i//GC(G),
(6.5.2) (φ, ψ) = — £ φ(χ)ψ(χ).
\g\x7g
Define the operators Ty\ у G G on C(G) by:
(6.5.3) for feC(G), Tyf(x)=f(x-y).
Observe that for у eG,Ty is unitary and, if the order of у is m, the
minimal polynomial minPT (z) is zm — 1, and has m simple roots, namely
the m'th roots of unity in T* (the multiplicative group of complex
numbers of absolute value 1). Observe also that the operators Ty commute,
and the subalgebra they generate is therefore self-adjoint.
Let φ G C(G) be a nontrivial common eigenvector of the operators
Ty, and denote by Xy the corresponding eigenvalue of Ty. We have
(6.5.4) <P(x~y) = hy<P(x)
so that if φ is nontrivial, then φ(χ) фО for all x, and since |Яу| = 1,
we have \φ(χ)\ constant for all xeG. We may replace φ by φ(0)_1 φ,
that is, assume with no loss of generality that φ(0) = 1, and hence
\φ(χ)\ = 1 for all x.
We now identify the eigenvalues Xy by checking (6.5.4) for x = y.
This gives 1 =Xy(p{y) and, since T_y = T~l, we obtain
(6.5.5) λ_γ = λ;ι=φ(γ),
and (6.5.4) becomes
(6.5.6) <P(x + y) = <P(x)<P(y)·
This shows that the φ is a character of G, that is, a homomorphism of
GintoT*.
By Theorem 6.3.6 there exists an orthonormal basis {7/}'2i f°r
C(G) consisting of common eigenvectors of the operators Ty.
By the preceding discussion, with φ being a prototype of the y-'s,
we see that, normalizing if necessary so that у;(0) = 1, every у is a
character of G.
6.5. Unitary and orthogonal operators
125
Lemma. Distinct characters on G are mutually orthogonal.
PROOF: If φ and ψ are distinct characters on G, then for every yEG,
(<p, ψ) = Σ 9(χ)Ψ(χ) = Σ <р(х-уЖх~у) = φΜ-1 ¥(у)(<р> ψ)
xeG xeG
and there exists у such that (p(y) Φ ψ(γ). <
Corollary. The set G= {y,} contains all the characters ofG.
PROOF: Since an orthonormal sequence in C(G) is linearly
independent, it cannot be a proper superset of the orthonormal basis {jj}·
There is no room for additional characters. <
Since {y } is an orthonormal basis for C(G), we obtain the Fourier
expansion for C(G).
Theorem. Every f £ C(G) has the following representation:
(6.5.7) /W=L(/,7)rW·
yeG
6.5.6 The set G is a multiplicative group, the product defined as the
pointwise multiplication inherited from C(G). It is commonly referred
to as the dual group of G.
For χ £ G the map χ : у н-> γ(χ) is a homomorphism of G into
T*, that is, a character on G. Since different elements χ £ G give rise
to different characters x, we obtain |G| = \G\ characters on G, which
means that we obtain all the characters on G. Direct checking that
x-\-y = xy completes the proof that the map χ н-> χ is an isomorphism
of G onto G, the dual of G.
EXERCISES FOR SECTION 6.5
ex6.5.1 Prove that the spectrum of a unitary operator is contained in the unit
circle {z : \z\ = 1}.
ex6.5.2 Prove that the set of rows of a unitary matrix is orthonormal.
126
6. Inner-Product Spaces
ехб.5.3 Let Τ £ &{Ж) be invertible and assume that \\TJ'\\ is uniformly
bounded for j e Z. Prove that Τ is similar to a unitary operator.
ex6.5.4 Show that if Τ G &{Ж) is self-adjoint and \\T\\ < 1, then there
exists a unitary operator U that commutes with T, such that Τ = ^(U + U*).
Hint: Recall that σ(Τ) С [-1,1]. For λ- e σ(Τ) write ζ. = Ay. + /vT^A?,
so that λ- = 91 ς. and |^| = 1. Define: Uv = Ιζ;·(ν,Uj)uy
ex6.5.5 Given A £ GL(n.C), define an inner product (·, -)A by
(6.5.8) (vx,v2)A = {AvvAv2).
Prove that the group of operators that preserve the inner product (·, ·)Α is the
conjugate subgroup, A-1 fy(n)A, of fy(n) in GL(n,C).
ex6.5.6 Let A G Ж{п\С) be diagonal with distinct eigenvalues. Show that
A — U~lBU is a Schur decomposition of A if and only if U is a permutation
matrix.
*ex6.5.7 Let G be a finite abelian group and G its dual. For every character
γ e G, denote y(G) = {y(jt) \xeG}. Verify the following statements.
a. y(G) is a cyclic subgroup of T*.
b. Let 70 G G be such that 70(G) is maximal, i.e., is not a proper subset of
7(G) for any 7 e G, and let x0 G G be an element of smallest order such that
70(;c0) is a generator for 70(G). Denote by X0 the subgroup of G generated
by jc0, and let G0 be the kernel of y0, i.e.,
G0 = {xeG:r0(x) = l}.
Then X0UG0 spans G, and X0nG0 = {0}, so that G = X0θG0.
c. If G0 is nontrivial then G is decomposable, i.e., is a direct sum of proper
subgroups. If G() is trivial then G is cyclic.
d. Every finite abelian group is a direct sum of cyclic groups.
This is essentially the basis theorem for finite abelian groups.
*ex6.5.8 Let G be a finite abelian group and G its dual group.
a. Prove that if G is cyclic then so is G.
b. If G = G{ ©G2. thenG = Gj ®G2.
c. The dual group G of a finite abelian group G is isomorphic to G.
*6.6. Positive definite operators
127
*6.6 Positive definite operators
6.6.1 An operator 5 is nonnegative definite, written 5 > 0, if it is self-
adjoint,5 and
(6.6.1) (5v,v) >0
for every ν G Ж. 5 is positive definite, written 5 > 0, if 5 > 0 and
(5v, v) = 0 implies ν = 0. We often drop the 'definite', and write simply
nonnegative or positive, as the case may be.
Lemma. A self-adjoint operator S is nonnegative definite if and only
if(j(S) С [Ο,οο), and positive definite if and only ifc(S) С (0,°о).
PROOF: Use the spectral decomposition 5 = £,· Α·7Γ·, where {A·} =
σ(Γ). We have (5v,v) = Σλ;·||π;·ν||2, and Σ||π;·ν||2 = ||ν||2. This is
clearly nonnegative for all ν G Ж if and only if A · > 0 for all j. It is
strictly positive for all ν φ 0 if and only if every A · is positive. <
6.6.2 Partial orders on the set of self-adjoint operators. Let Τ and
5 be self-adjoint operators. The notions of positivity and nonnegativity
define partial orders, ">" and ">", on the set of self-adjoint operators
on Ж. We write Τ > S if Τ -S > 0, and Τ > S if T-S > 0.
Proposition. Let Τ and S be self-adjoint operators on Jff, with Τ >S.
Let o(T) = {A-}y=1 and (7(5) = {Д;И=р both arranged in nonde-
creasing order. Then A > μ for all j.
PROOF: Use the minmax principle, exercise ex6.3.4:
A, = min max (7V,v) > min max (5v, v) = u,.
J dimW=j ve^,||v|| = l dimW=j veW,\\v\\ = \ J ^
Remark: The condition "A > μ for 7 = 1,... ,n" is necessary but,
even if Τ and 5 commute, not sufficient to prove that Τ > S, except
5The assumption that S is self-adjoint is supefluous—it follows from (6.6.1). See
8.2.3.
128
6. Inner-Product Spaces
when η = 1. Consider for example the operators Γ, 5: C2 —> C2 defined
as left multiplication by the matrices
/ly
2 0
0 4
and A5
3 0
0 1
The eigenvalues for Τ (in nondecreasing order) are {4,2}, for 5 they
are {3,1}, yet for Τ — 5 they are {3,-1}.
*6.7 Polar decomposition
6.7.1 Lemma. A nonnegative operator S on Ж has a unique nonneg-
ative square root.
PROOF: We use the spectral theorem to write 5 =ΣλΕσ/^λπλ, where
πλ is the orthogonal projection on Ж^ Define \/~S =Σ\ίλπλ, where
we take the nonnegative square roots of the (nonnegative) A's. Then
л/5, being a linear combination with real coefficients of self-adjoint
projections, is self-adjoint, and (л/5)2 = 5.
To show the uniqueness, let Τ be nonnegative with T2 = 5. Then
Τ and 5 commute, and Τ preserves all the eigenspaces Жх of 5.
On ker(5), if 0 £ <r(5), then T2 = 0 and, since Τ is self-adjoint,
Τ = 0. On each Ж^ for λ > 0, we have 5 = λΐχ (the identity operator
on Ж^) so that Г = \/Χ/λ, with л/Я > 0,/λ positive, and /^ = Ι χ. The
eigenvalues of J^ are ± 1, and the positivity of J χ implies that they are
all 1, so Jλ — Ιλ and Τ = л/5. <
6.7.2 Lemma. Let Ж- С ^, j = 1,2, be isomorphic subspaces. Let
f/j fee a linear isometry Ж] —> J^. 7%еи ί/геге are unitary operators
U on Ж that extend U].
PROOF: Define U = U] опЖх, while on Ж^ define U as an arbitrary
linear isometry onto Ж^~ (which has the same dimension). Extend by
linearity to all of Ж. л
6.7.3 Lemma. Let A,fiG ^(Ж); then there exists a unitary
operator U such that В = AU if and only if \\Av\\ = ||Bv|| for all ν £ Ж.
Furthermore, U is unique if and only //>ange(A) = Ж.
*6.7. Polar decomposition
129
PROOF: If U is unitary and В = f/A, then ||Bv|| = \\UAv\\ = ||Av||, for
all уеЖ
Next, we note that ker(B) = ker(A). So, if range(A) = Ж, then
range(Z?) = Ж. If {u],... ,и„} is an orthonormal basis of Ж, then
{Ам1,...,Ами} and {βίΐρ.,.,β^} are both bases of Ж. The map
i/ : Am h-> Bm . extends to a linear isometry on all of Ж, and U = BA~l
is unique.
If ker(A) φ {0}, then let {м1,..., un} be an orthonormal basis of
Ж such that {wp...,ww} is a basis for ker(A) = ker(B). The sub-
space range(A) is spanned by {Au,}^+1 and range(Z?) is spanned by
{Вы ;}JJj+1. The map f/j : Am h-> #w. extends by linearity to an
isometry of range(A) onto range(#). Now apply Lemma 6.7.2, and
remember that U = υλ on the range of A. Since U can be defined as an
arbitrary linear isometry on the orthogonal complement of range(A),
it follows that U is not unique in this case. <
6.7.4 We observed in 6.3.1, that for any Τ e Jf(Jff), the operators
Г*7 and 7Ύ* are self-adjoint. Remember, however, that in general
τ*τ φ TT* (uniess τ is normal).
For any ν e Ж
(7*7v,v) = (Γν,Γν) = ||Γν||2 and (7Γ*ν,ν) = ||Γ*ν||2,
so that Γ* Γ and 7Ύ* are both nonnegative, and by Lemma 6.7.1, each
has a nonnegative square root. We shall use the notation
(6.7.1) \T\ = Vf*f.
Theorem (Polar decomposition). Every operator Τ e Jf(Jf?)
admits a representation as a product
(6.7.2) T = U\T\,
where U is unitary, and \T\ is nonnegative.
PROOF: Observe that
||Γν||2 = (Γν,Γν) = {ΓΤν,ν) = (\Τ\\ν) = (|Γ|ν, |Γ|ν) = |||Γ|ν||2.
130
6. Inner-Product Spaces
The theorem now follows from Lemma 6.7.3, with A
\T\<mdB = T.
When dim^f = 1, this reduces to the polar decomposition of a
complex number: ζ = \z\e1^. Polar decomposition should not be
confused with the polarization formula, (6.1.15).
Remark: Starting with Γ* instead of Τ we obtain a unitary operator
Ux such that
(6.7.3) 7* = C/, | Γ* | = и{л/ТТ*.
Taking adjoints, we also obtain a representation of the form
(6.7.4)
T = \T*\U-
ι ·
Notice, however, that in general |7*| and \T\ may be different. For
example, let Τ be the map on C2 defined by Tvx = v2 and Tv2 = 0,
where Vj, v2 are the elements of an orthonormal basis. Then T*Vj = 0
and 7*v2 = Vj so that T*T, as well as its nonnegative square root \T\, is
the orthogonal projection onto the line of the scalar multiples of Vj. On
the other hand, 7Ύ*, as well as |7*|, is the orthogonal projection onto
the multiples of v2. The matrices corresponding to the basis {vj, v2}
are:
T =
"o o"
1 0
5
and £/ = £/,=
T*
"ο ι"
1 0
"о Г
0 0
ГГ1*гр
1 0
о о
rrirp*
о о
0 1
6.7.5 Given Τ Ε ££{Ж), the eigenvalues {μΙ,... ,μ„} of the non-
negative operator \T\ = y/T*T, are called the singular values of 7\
Let {iij,..., un] be corresponding orthonormal eigenvectors, and
denote ν · = С/и ·, where U is a unitary operator satisfying (6.7.2).
Being the image of an orthonormal sequence by a unitary operator,
{vj,..., vn] is orthonormal.
For all ν Ε Ж, ν = ]C(v,n -)и ·, |Γ|ν = Σμ;·(ν,ιι -)иу·, and
(6.7.5)
Γν = £μ·(ν,ΐί.)ν..
*6.7. Polar decomposition
131
This is sometimes written6 as
(6.7.6) T = yL^juj®vj'
EXERCISES FOR SECTION 6.7
ex6.7.1 Let {wx,..., w„} be an orthonormal basis for Ж and let Τ be the
(weighted) shift operator on {wx,..., wn} defined by: Twn = 0 and for j < n,
Tw- = in — j)w-+x. Describe U and \T\ in (6.7.2), as well as |7*| and Ux in
(6.7.3).
ex6.7.2 An operator Τ is said to be bounded below by c, on a subspace
rcJ, written: Τ > с on У, if ||Γν|| > c||v|| for every vGl
Assume now that {ux,..., un} and {vj,..., vn} are orthonormal sequences,
μλ > μ2 > ' · > Vn > 0, and Τ e а?(Ж) is defined by Τ = £μ;·Μ;· ® vy. Show
that μ = max{c : there exists a y-dimensional subspace on which Τ > с }.
ехб.7.3 The following is another way to obtain (6.7.6) in a somewhat more
general context. It is often referred to as the singular value decomposition of
T.
Let Ж and Ж be inner-product spaces, and Τ £ ^{Ж, Ж) of rank
r > 0. Write μχ = || Γ || = max,,,, 1 ||Γν||. Let vx £ Ж be a unit vector such
that ||rvj || = μ1? and write zx = μ^Γν^ Observe that zx £ ^ is a unit
vector.
Prove that if r = 1, then 7 = μ1 Vj 0 Zj, that is, Τ ν = μχ (ν, Vj )zj, for all
Assuming r > 1, let Tx = Τ f ,±, the restriction of Γ to the orthogonal
complement of the span of vx; let πχ be the orthogonal projection of Ж onto
spanfzj]1. Write μ2 = \\κχΤχ ||, let v2 e spanfvj]1- be a unit vector such that
H^i^iSlI = A*2 anc* WI"ite^2 = ^2]π\
discontinue in this way to define μ , ν · and 7 · recursively for у < г, so that
a. The sets {v ·} <r and {ζΛι<κ are orthonormal in Ж, respectively Ж.
b. For m<r, μιη is defined by μ™ = \\жт_хТт_х ||, where Гт-1 denotes the
restriction of Τ to spanfvj,..., vm_1]±, and ^m-1 is the orthogonal projection
of JTontospan^p...,^!]-1.
с vm is such that ||^w_1 Гт-1 vm|| = μm, and zm = μ7ηΧππι-\Τπι-\ν^
6See *4.2.5.
132
6. Inner-Product Spaces
Prove that with these choices,
(6.7.7) T = ^jVj®Zj.
*6.8 Contractions and unitary dilations
6.8.1 A contraction on Ж is an operator whose norm is < 1, that is,
such that ||7V|| < || v|| for every ν G Ж. Unitary operators are
contractions; orthogonal projections are contractions; the adjoint of a
contraction is a contraction; the product of contractions is a contraction.
If 7 G ££{Ж) is a contraction, then both 7*7 and 77* are positive
contractions and so are / — 77* and / — 7*7, and we set
(6.8.1) A = (/-77*)5 and Я = (/-ГТ)г.
Lemma. Let Τ be α contraction and let A and В be defined by (6.8.1).
Then AT = ТВ.
PROOF: First note that A2k = (1 — 77*)* is a linear combination of
powers of TT*, while Blk = (1 - T*T)k is the same combination in
which each power of 77* is replaced by the same power of 7*7. Since
7(7*7)* = (77*)*7 for all к > 0, it follows that for every positive
integer k, we have
(6.8.2) A2kT = TB2k and hence P(A2)T = TP(B2)
for every polynomial P.
Since σ(7*7) U σ(77*) is a finite subset of [0,1], there is a
polynomial P(x) = Y^akxk such that P(x2) = χ on σ(7*7) U σ(77*), and
hence A = P(A2) and В = P{B2). It now follows from (6.8.2) that
A7 = P(A2)7 = TP{B2) = ТВ. <
6.8.2 Let <Щ be a subspace of an inner-product space Jff, and
denote by πλ the orthogonal projection of Ж onto Жх.
Definition: A unitary operator U on Ж is a unitary dilation of a
contraction 7 G Jf(Jff[) if 7 = (πχυ)^.
Theorem, /fdim^f = 2dimJiflf then every contraction in Jzf(^)
has a unitary dilation in Ж.
*6.8. Contractions and unitary dilations
133
PROOF: Write Ж = Жх ®Ж2, with Ж2 _L Жх. Let ν = {ν,,..., ν„}
be an orthonormal basis for Жх and u = {uv...,un} an orthonormal
basis for Ж2. Then vu = {vx,..., v„, ux,..., un} is an orthonormal
basis for Ж.
Let Τ be a contraction on ^, and define A and Б by (6.8.1).
Let U be the operator on Ж whose matrix relative to the basis vu
Τ A
', where the operator names stand for the η χ η matrices
is
-Β Τ
corresponding to them in
л; C) relative to the basis ν of Жх.
The matrix of %x relative to vu is
7 0
0 0
so πχϋ =
г
Τ Α
О О
, and
Τ Α"
-Β Τ*
Τ*
Α
-Β
Τ
■ ττ*+Α2
Τ*А-ВТ*
AT-ТВ
Β2 + Τ*Τ
hence, (n^U)^ = Т.
To show that U is a unitary dilation of 7\ it remains to show that
U is unitary, i.e., we need to show that U* = U~]. We have
(6.8.3) UW
The terms on the main diagonal of the product reduce to /, and the
terms off the diagonal are adjoints of each other. Since AT — TB = 0,
by Lemma 6.8.1, it follows that UU*=I. <
6.8.3 One can state the theorem in terms of matrices rather than
operators.
Definition: A matrix in J^{n\C) is contractive if it is the matrix
of a contraction relative to an orthonormal basis.
Theorem. A matrix in Ж{п\С) is contractive if and only if it is the
top left quarter of some 2n χ 2η unitary matrix.
EXERCISES FOR SECTION 6.8
ex6.8.1 Check that the "only if" part of Theorem 6.8.3 is done in the proof
of Theorem 6.8.2, and prove the "if" part of the theorem.
ex6.8.2 An m χ / complex matrix A is contractive if the map it defines in
.Sf (Cz,Cm) (with the standard bases) has norm < 1.
134
6. Inner-Product Spaces
An η χ к submatrix of an m χ / matrix Α, η < m,k < /, is a matrix obtained
from Λ by deleting m — nof its rows and l —к of its columns. Prove that if A
is contractive, then every nxk submatrix of Л is contractive.
ex6.8.3 Let Μ be an m χ m unitary dilation of the contraction Τ = 0 £ Ж (η).
Prove that m>2n.
Chapter 7
Structure Theorems
7.1 Reducing subspaces
7.1.1 Let (У,Т) be a linear system. A subspace Ψλ С Ψ reduces
Τ if it is 7-invariant and has a 7-invariant complement, that is, there
exists a Γ-invariant subspace У2 such ^^ У = У]®У2-
Example: If Ψ = Ж is an inner-product space and Τ e ££{Ж)
is self-adjoint, then every 7-invariant subspace is reducing since its
orthogonal complement is also ^-invariant. See Proposition 6.3.1 b
A system (У,Т) that admits no reducing subspaces is called
irreducible. We say also that Τ is irreducible on Ψ. А Γ-invariant subspace
W is irreducible if Ίψ is irreducible.
Theorem. Every system (Ψ, Τ) can be decomposed into a direct sum
of irreducible systems.
PROOF: Use induction on η = dim У'. If η = 1, the system is trivially
irreducible. Assume the validity of the statement for η < N and let
(V,T) be of dimension N. If (У,Т) is irreducible, the
decomposition is trivial. If (У, Т) is reducible, let Ψ = Ψλ Θ У2 be a n°ntrivial
decomposition with 7-invariant Ψ-. Then άιναψ- < Ν; hence each of
the systems {Ψ·,Τψ) is completely decomposable, Ψ- = ®кУ;к with
every Ψ. k 7-invariant, and Ψ = 0 . k У. k is a complete decomposition
of(r,r). <
7.1.2 The reason for our interest in reducing subspaces and
decompositions is that the operator can be analyzed separately on each direct
summand of a decomposition. This reduces the analysis of the
structure of a general linear system to that of the structure of irreducible
135
136
7. Structure Theorems
systems, and we shall see in subsections 7.1.6, 7.3.3, and in section
*7.5 that irreducible systems have a relatively simple structure:
a. If (У,Т) is irreducible, then the minimal polynomial minPT and
the characteristic polynomial χτ coincide, and are equal to a power,
Фт, of an irreducible polynomial Φ G F[jc].
b. Irreducible systems are cyclic.
As we saw in 5.3.2, if (Ψ, T) is cyclic of dimension к and ν e У
is a cyclic vector, then the matrix A7, relative to the basis {Th}kr}^ is
the companion matrix of minPT.
7.1.3 A direct-sum decomposition of У into 7-invariant subspaces
allows us to find a basis for Ψ, with respect to which the matrix of Τ
is particularly simple. This can be seen as follows.
Let dim^ = n, and assume that ψ = ψχ®Ψ2, where ΨΧ and Ψ2
are 7-invariant. Let {vp..., vn} is a basis for Ψ such that the first к
elements are a basis for ΨΧ while the last I = n — k elements are a basis
forr2.
The entries a{ ■ of the matrix AT of Τ relative to this basis are zero
unless / and j are both < k, or are both > k. This means that AT
consists of two square blocks centered on the main diagonal. The first
is the A: x A: matrix of Τψ (relative to the basis {vp... ,ν^}), and the
second is the / χ / matrix of Ту (relative to {v^+1,..., v„}).
More generally, assume that Ψ = ®)=\ %, with ^-invariant
components Ψ- of dimension k·. Let ν = {и- ·} .J_ be a basis for Ψ-, then
ν = (J \j; is a basis for У, and we enumerate its elements by listing in
their order the elements of vp followed by those of v2, etc.
If A is the matrix of Ty relative to the basis ν , for j = 1,... ,s,
j j j
then the matrix A7, of Γ relative to the basis v, is the diagonal sum1
of the matrices A ·. That is, AT consists of the s matrices AV...,AS
1 Also called the direct sum ofA-,j= 1,..., s.
7.1. Reducing subspaces
137
centered along the diagonal, and has zero entries everywhere else:
(7.1.1) AT =
Άχ 0 0 0
0 A2 0 0
0 0 A, 0
0 0 0
7.1.4 Theorem 7.1.1 tells us that every linear system is a direct sum
of irreducible subsystems, but does not tell us how to find them. Our
program now is to show that the decomposition runs parallel to the
decomposition of minPT as a product of powers of irreducible
polynomials, minPT = ЦФт'. This leads to the canonical prime-power
decomposition of Ψ as a direct sum of the corresponding nullspaces,
кег(Ф^))·
Generally speaking, the kernel of an operator Τ is not reducing.
The following theorem, gives a complete characterization of operators
whose kernels are reducing.
Theorem. Let (Ψ\T) be finite-dimensional ker(T) and range(T) are
reducing subspaces if and only ifker(T) Π range(T) = {0}, in which
case, the only Τ-invariant complement ofker(T) is range(T), and vice
versa.
PROOF: If ker(7) Π range(7) = {0}, then the sum ker(7) + range(7)
is a direct sum, so the dimension of the sum is the sum of the
dimensions. By the rank and nullity theorem (see subsection 2.5.1), we have
dim(ker(7)erange(7)) = ν(Γ) +ρ(Τ) = dim^,
so^ = ker(T)©range(T). Since ker(T) and range(T) are ^-invariant,
it follows that both spaces are reducing.
If Ψχ is a 7-invariant subspace such that Ψ = ker(T) Θ V^ then
range(7) = ТУХ С Ψν On the other hand,
dim^ =dimr-dim(ker(7)) = dim(range(7)),
so Ψχ = range(7), and ker(7)Пrange(7) = ker(7)Г)УХ= {0}.
7. Structure Theorems
Finally, if 7^ is a ^-invariant subspace and Ψ = Ψ2 Θ range(T),
then ΤΎ2 С ^nrange(7) = {0}, so Ψ2 С кег(Г). One more appeal to
the rank and nullity theorem shows that dim ^ = V(T)> so ^2 = кег(Г)
and кег(Г) Π range(r) = {0}. ^
Remark: The uniqueness of the Γ-invariant complement for кег(Г)
and range(r) is special. For general reducing subspaces there is no
such uniqueness, as you can see by considering the identity operator,
for which every subspace is reducing.
Corollary. кег(Г) reduces Τ if and only if
кег(Г2) = кег(Г).
PROOF: For every Τ e &(Ψ) we have кег(Г2) D кег(Г). The
inclusion is proper if and only if there exist vectors ν such that Τ ν φ 0 but
Γ2ν = 0, which amounts to Tv e range(T) Π кег(Г). м
7.1.5 Basic decomposition. Let (У,Т) be an η-dimensional
system. In the theorem below, we describe the fundamental relationship
between factorizations of minPT and decompositions of Ψ into direct
sums of Γ-invariant subspaces. We begin with a proposition that
illuminates the basic ideas.
Proposition. If the polynomials Ρχ and P2 are relatively prime, then
(7.1.2) кег(Р1(Г))Пкег(Р2(Г)) = {0}.
If also Px (T)P2(T) = 0, then У = ker(P, (Τ)) Θ кег(Р2(Г)), and the
corresponding projections are polynomials in T.
PROOF: Given that P] and P2 are relatively prime, there exist, by
Corollary A.6.2, polynomials ql, q2 such that qxP{ -\-q2P2 = 1.
Substituting Τ for the variable we have
(7.1.3) q,{T)P{{T)+q2{T)P2{T)=I.
If ν e ker(P, (Г)) П кег(Р2(Г)), that is, Ρ, (7> = P2(T)v = 0, we
have ν = qx{T)Px (T)v + q2(T)P2(T)v = 0. This proves (7.1.2), which
implies, in particular, that dimker^, (Γ)) + dimker(P2(T)) < n.
7.1. Reducing subspaces
139
If Pl(T)P2(T) = 0, then range(P,(r)) С кег(Р2(Г)) and likewise,
range(P2(r)) С ker(P{(T)). By the rank and nullity theorem
η = dimker(P, (Γ)) +dim ranged (Γ))
<dimker(P1(r))+dimker(P2(r))<n.
It follows that dimkerfT^r)) + dimker(P2(r)) = n, which implies
that ker(Pj (Γ)) Θ кег(Р2(Г)) is all of Ψ.
Equation (7.1.3) implies that
<f>2{T)=qx{T)Px{T)=I-q2{T)P2{T)
is the identity on кег(Р2(Г)) and zero on ker (P^T)), i.e., φ2(Τ) is the
projection onto кег(Р2(Г)) along ker(P{(T)).
Similarly, φλ(Τ) = q2(T)P2(T) is the projection onto кег(Р,(Г))
along кег(Р2(Г)). <
Remark: The proof also shows that, if P] P2 = 0, then
range(P1(r)) = ker(P2(r)) and P2(T) = ker (Рг(Т)).
The general relationship between factorizations of minPT and
decompositions of [Ψ, Τ) is described by the following theorem.
Theorem. For every factorization, minPT = Π;=ιΛ> into pairwise
relatively prime factors, we have a direct sum decomposition of Ψ
into Ί-invariant subspaces,
ι
(7.1.5) Г = 0кег(р.(г)),
7=1
and the projection onto each direct summand is a polynomial in T.
PROOF: We use induction on the number of factors, /.
For / = 2, the theorem follows immediately from the preceding
proposition, so assume that / > 2, and that the statement is true for
/ — 1. Write Q = T[lj=2Pj, then by the same proposition we have
Г = кег(Р1(Г))екег(б(Г)).
140
7. Structure Theorems
Applying the induction hypothesis to (ker(Q(T)),Tker{Q{T))) gives
кег(е(Г)) = 0кег(Р;(Г)),
7=2
and the stated decomposition of У follows.
For 1 < ; < /, let Q- = U¥jP^ then gcd(P;, Qj) = 1, so by
Corollary A.6.2 there exist polynomials q- and p- such that q-P-+p -Q- = 1.
Evaluating this identity at Τ gives qj(T)P-(T) +pAT)Q-(T) = /, from
which it follows, as it did in Proposition 7.1.5, that the polynomial
(7.1.6) φ.(Γ)=/7.(Γ)ρ;(Γ)=/-^(Γ)Ρ.(Γ)
is the projection onto кег(РДГ)) along 0;.,. кег(РДГ)). ^
Notice that the projections φλΤ) all commute, since they are all
polynomials in Τ. Also, they are mutually orthogonal in pairs, in the
sense that φί(Τ)φ](Τ) = 0 when ; ф i.
7.1.6 The canonical prime-power decomposition. Applying
Theorem 7.1.5 to the prime-power factorization of minPT,
ттРт = Пф?>
where the Φ ,'s are distinct irreducible (prime) polynomials in F[jc], and
m their respective multiplicities, we obtain the canonical prime-power
decomposition of (У, T)\
к
(7.1.7) Г = 0кег(ф^(г)).
The spaces кег(ФШ; (Г)) are called the primary components of [Ψ, Τ).
The projection onto кег(ФШ;(Г)) is a polynomial in Г, as given
above in (7.1.6),
(7.1.8) φ;.(Γ)=/-^.(Γ)Φ^(Γ)=ρ;.(Γ)ΠΦΓ'(Γ),
¥J
2See Theorem A.6.3.
7.1. Reducing subspaces
141
where q- and ρ are polynomials satisfying q Φ. ' + /?,-П^/Фт'" — 1·
If Ψ С Ψ is Γ-invariant, then the subspaces
ф;.(Г)Ж = >ГПкег(Ф^(Г)),
are Γ-invariant and we have a decomposition
(7.1.9) Г = 0(р.(г)Г
7=1
Proposition. Γ/ze Ί-invariant subspace W is reducing if and only if
φ:(T)W is a reducing subspace of кег(ФШ](Т)) for every j.
PROOF: If W is reducing and ^ is a Γ-invariant complement, then
кег(Ф^(Г)) = <f>j{T)r = q>j{T)W®q>j{T)<%,
and both components are Γ-invariant. Conversely, if ^. is Γ-invariant
and кег(Фт;(Г)) = φ^Ύ)Ψ Θ ^, then У/ = 0 <&. is an invariant
complement to W. <
7.1.7 By the Cay ley-Hamilton theorem and Corollary 5.3.4, the prime
factors of χτ are those of minPT, with at least the same multiplicities,
that is:
(7.1.10) χτ = γ[φΐ, with Sj>mr
The minimal polynomial of the restriction of Γ to кег(ФШ;(Г))
is ФШ] and its characteristic polynomial is Φ5/. The dimension of
кег(Фт;(Г))18^её(Ф7·).
./ j j
7.1.8 When the underlying field F is algebraically closed, and in
particular when F = C, every irreducible polynomial in ¥[x] is linear and
every polynomial is a product of linear factors, see A.6.6.
Recall that the spectrum of Г is the set σ(Γ) = {A ·} of zeros of
χτ or, equivalently, of minPT. For systems over an algebraically closed
field the prime-power factorization of minPT has the form
minPT= Ц (x-A)mW
λεσ(Τ)
142
7. Structure Theorems
where m(A) is the multiplicity of λ in minPT.
The space Ψχ = кег((Г — A)w^) is called the nilspace, or the
generalized eigenspace, of A. The canonical decomposition of (У ,T) is
given by:
(7.1.11) Г= 0 Гя.
Aea(r)
The same decomposition obtains more generally when all the prime
factors of minPT are linear (even when F is not algebraically closed).
EXERCISES FOR SECTION 7.1
ex7.1.1 Let Τ e -Sf (У), k > 0 and integer. Prove that кег(Г^) reduces Τ if
and only if кег(Г*+1) = кег(Г^).
Hint: Both кег(Г*) and range(7^) are ^-invariant.
ex7.1.2 Let Τ e &(У), and У = W Θ Ψ with both summands ^-invariant.
Let π be the projection onto <%£ along Ψ. Prove that η commutes with T.
ex7.1.3 Prove that if (У ,T) is irreducible, then its minimal polynomial is
"prime power", that is, minPT = Фт with Φ irreducible and m > 1.
ex7.1.4 If У- = кег(ФШ; (Τ)) is a primary component of (У, T), the minimal
polynomial of Ty is Φ.;.
ex7.1.5 Assume that Τ is nonderogatory, i.e., %T = minPT.
a. Prove that if g e ¥[x] divides minPT, then dimker(g(r)) = degg.
b. Prove that iffe¥[x], then
(7.1.12) dimker(/(r)) = deggcd(minPT,/).
Hint: Write f(T) = g(T)h(T) where g = gcd(/, minPT).
7.2 Semisimple systems
7.2.1 Definition: The system (У,Т) is semisimple if every T-
invariant subspace of У is reducing.
Self-adjoint operators on inner-product spaces give the standard
examples of semisimple systems.
7.2. Semisimple systems
143
Theorem. The system (У,Т) is semisimple if and only j/minPT is
square-free (that is, the multiplicities m · of the factors in the canonical
factorization minPT = П*Ш; are °^ Ц
PROOF: Proposition 7.1.6 reduces the general case to that in which
minPT is Фт with Φ irreducible.
If m > 1, then Φ(Γ) is not invertible, and hence the invariant
subspace кег(Ф(Г)) is nontrivial nor is it all of Ψ. кег(Ф(Г)2) is
strictly bigger than кег(Ф(Г)) and, by Corollary 7.1.4, кег(Ф(Г)) is
not Ф(Г)-геаиси^, and hence not Γ-reducing.
If m = 1, we observe first that minPT v = Φ for every nonzero
vector ν e У. This is true since minPTv divides Φ and Φ is prime. It
follows that the dimension of span[r, v] is equal to the degree d of Φ,
and hence: every nontrivial Τ-invariant subspace has dimension > d.
Let Fcfbea proper Γ-invariant subspace, and Vj еУ\Ж. The
subspace span[r, vj Π W is Γ-invariant and is properly contained in
span[r, vj (since Vj 0 W), so its dimension is less than d. This means
thatspan^vjRyT = {0}, so span[r,Ψ U {vj] = Ψ0span[r,v,].
Let Г1=Г0 spa η [Γ, νλ].Ψλ is Γ-invariant and if Ψλ φ Г, then
by the preceding argument, if v2 G Ψ \ Wx, then
span[r,>T1U{v2}] = >respan[r,v1]espan[r,v2].
After к = (dim ψ — dim W)/d repetitions of this process, we have
(7.2.1) r = ^e05=1span[r,vy.],
where 0j span[r,v·] is clearly Γ-invariant, so Ψ is reducing. <
Remark: If we start with Ψ = {0}, the decomposition (7.2.1) of
(У,Т) is a direct sum of cyclic subsystems.
7.2.2 If F is algebraically closed, then the irreducible polynomials in
F[jc] are linear. In other words, the prime factors Φ of minPT have the
form χ — λ ·, with A, G <?(T).
144
7. Structure Theorems
By Theorem 7.2.1, if (У, T) is semisimple, then
(7.2.2) minPT(*) = Ц (x-Xj)
λ;βσ(Τ)
and the canonical prime-power decomposition has the form
(7.2.3) Г = 0кег(Г-*у) = 0ф;.(Г)Г,
where φ AT) are the projections defined in (7.1.8). The restriction of
Τ to the eigenspace кег(Г — A·) = φΑΤ)Ψ is just multiplication by
A , so that
(7.2.4) Τ = ΣλιΨ]{Τ),
and for every polynomial P,
(7.2.5) Р(Г) = £Р(Я.)<Р;(Г).
A union of bases of the respective eigenspaces кег(Г — A ·) is a basis
for У whose elements are all eigenvectors, and the matrix of Τ relative
to this basis is diagonal. This proves the following theorem.
Theorem. A semisimple operator on a vector space over an
algebraically closed field is diagonalizable.
7.2.3 Definition: An algebra 88 с 3?(У) is semisimple if every
Τ G 88 is semisimple. SB is diagonalizable if Ψ has a basis whose
elements are all common eigenvectors of all the operators in S3.
Equivalently, SB is diagonalizable if Ψ has a basis relative to which
the matrices of all the elements of SB are diagonal.
Theorem. Assume that ¥ is algebraically closed and let S3 С ££{Ψ)
be a commutative semisimple subalgebra. Then
a. 88 is diagonalizable.
b. 88 is linearly spanned by the projections it contains.
с 88 is singly generated, that is: there are elements Τ G 88 such that
8$ = 8*(T) = {P(T) : Ρ e¥[x}}.
7.2. Semisimple systems
145
Proof:
a. Let Τ G SB be such that σ(Γ) is maximal (in terms of the number
of its elements) and let
(7.2.6) Г= 0 Гя
Aea(r)
be the canonical decomposition of Ψ, where >^ = кег(Г — A).
Each >^ is invariant under every operator 5 in ^, and the
restrictions Ξλ = Sy are semisimple. We claim that Ξλ acts as a scalar on
n·
To see this, first note that if μ G σ(5λ), then Я +αμ G σ(Γ + я5),
for any a G F. Suppose now that for some S e & and some A G σ(Γ),
the spectrum of Ξλ has more than one element. Since F is infinite,3
it follows from exercise ex7.2.1 that we can find an a G F such that
σ(Τ + aS) has more elements than o(T). This contradicts the choice
of Γ.
Thus, for all 5 G SB and A G σ(Γ), 5Я acts as scalar, so every
element of Ψχ is an eigenvector for every operator in SB. Take a basis
in each Ψχ\ their union is the required basis of Ψ.
b. Let Τ G ^ be an operator with maximal spectrum, as above, and
let Ψ = ®Υλ be the corresponding canonical decomposition into the
eigenspaces of T. By Theorem 7.1.5, the projection onto Ψχ is a
polynomial in Γ, φλ(Τ), so these projections are all in ^.
By part a, for any 5 G ^ and A G ст(Г), there is a scalar μλ such
that Sv = μλν for all ν G ^. It follows that
(7.2.7) 5= £ Мя-ФяСО-
ΑΕσ(Γ)
c. The right-hand side of (7.2.7) is a polynomial in Г. <
Remarks: The projections {фя(Г) : A G <т(Г)} (for Τ with maximal
spectrum) form a basis for ^, see exercise ex7.2.4.
'Recall that an algebraically closed field must be infinite.
146
7. Structure Theorems
Different approaches to proving parts b and с are given in exercises
ex7.2.5 and ex7.2.6, respectively.
EXERCISES FOR SECTION 7.2
ex7.2.1 Let {λλ, μι),..., (Ял, μ^) be к distinct points in F2. Show that there
are only finitely many a e¥ such that Я- + αμϊ = Я + αμ , for some pair
\<i<j<k.
ex7.2.2 If Τ is diagonalizable (У has a basis consisting of eigenvectors of
T) then (У, T) is semisimple.
ex7.2.3 Let У = 0 Жbe an arbitrary direct sum decomposition of У, and
let 7Г be the corresponding projections. Let Τ = Σ J71
ία. Prove that Γ is semisimple.
e. Exhibit polynomials P: such that 7Γ = Ρ AT).
ex7.2.4 Show that the projections {φλ(Τ) : Я £ ст(Г)} that appear in the
proof of part b of Theorem 7.2.3 are linearly independent.
Hint: See the remark that follows the proof of Theorem 7.1.5.
ex7.2.5 Let 38 с ££{Ύ) be a commutative subalgebra. For projections in 38,
write π{ < π2 if κχκ2 = πν
a. Prove that this defines a partial order on the set of projections in 38.
Hint: Check that:
i. if π{ < π2 and π2 < π{ then πχ = π2\
ίί. if πj < тг2 and π2 < π3, then πχ <л3.
b. A projection π Φ 0 is minimal if 0 ^ л^ < π implies π{ = тт. Prove that
every projection in 38 is the sum of minimal projections.
c. Prove that if 38 is semisimple, then the set of minimal projections is a
basis for 38.
ex7.2.6 Prove part с of Theorem 7.2.3, by using the Lagrange interpolation
theorem, see exercise ex3.1.8, part с
Hint: Using the notation in the proof of the theorem, show that S = P(T),
where Ρ e ¥[x] is the polynomial satisfying Ρ (λ) = μλ, for Я е о (Т).
ex7.2.7 Let У = У{ Θ У0 and let 38 С .Sf (У) be the set of all the operators
S such that S1\ С У0, and S^ = {0}. Prove that 38 is a commutative
subalgebra of if (У) and that dim^ = dim У0 dim ^. When is «^ semisimple?
7.3. Nilpotent operators
147
ex7.2.8 Let 38 be the subset of J%(2\ R) of the matrices of the form
a b
—b a
Prove that 38 is an algebra over R, and is in fact a field isomorphic to C.
ex7.2.9 Let Ψ be an «-dimensional real vector space, and Τ e ££(Ύ) an
operator such that minPT(x) = Q(x) = x2 + bx + c, with b2 < Ac.
Prove that η is even and, for an appropriate basis in У, the matrix AT
α β
consists of η/2 copies of
-β α
along the diagonal, where a = b/2 and
/} = л/4с-62.
In what sense is (У ,Т) "isomorphic" to (СИ//2,А/)? (/ the identity on
C"/2).
7.3 Nilpotent operators
The canonical prime-power decomposition (see subsection 7.1.6)
reduces every linear system to a direct sum of systems whose minimal
polynomials are each a power of an irreducible polynomial Ф.
If F is algebraically closed, and in particular if F = C, the
irreducible polynomials are linear, equal to (χ — λ) for some scalar
Α(Ξσ(Γ).
We consider here the (very important) special case in which all the
irreducible factors Φ are linear. In other words, we focus on irreducible
systems whose minimal polynomial has the form minPT = (x — A)m.
We discuss the general case in section *7.5.
7.3.1 If minPT = (x- A)m, then minP(T,Я) = xm. As Τ and Τ - λΐ
have the same invariant subspaces and the structure of Τ — λΐ tells
everything about that of Γ, we focus on the case λ = 0.
Definition: An operator Τ e 3?(У) is nilpotent if Tk = 0 for some
positive integer k. The smallest positive к for which Tk = 0 is called
the height of Τ and denoted height[T]. We refer to a linear system
(Ψ, T) as nilpotent of height A: if Γ is nilpotent and heightfr] = k.
If Tk = 0, minPT divides xk, hence it is a power of χ
If height[r] = k, then Γ*"1 φ 0 and minPT(x) = **. In other words,
Τ is nilpotent of height к if and only if minPT(x) = xk.
For every ν еУ, minPT v(jc) = xl for an appropriate /. The height
of ν (under the action of Г), denoted height [v], is the degree of minPT v,
148
7. Structure Theorems
that is, the smallest integer / such that Tlv = 0; it is the height of 7L·
where W = span[r, v]. Clearly height[T] = maxvGr height[v].
Since for ν φ 0, height[7V] = heightfv] — 1, elements of maximal
height are not in range(Γ).
7.3.2 Definition. A k-shift is a cyclic nilpotent system (У,Т) of
height k. A standard shift is a &-shift for some k, that is, a cyclic
nilpotent system.
If {У,Т} is a &-shift, and v0 £ Ψ is a vector of height k, then
ν = {T^v0}kZQ is a basis for У, and the action of Τ is to map each
basis element, except for the last, to the next one, and map the last
basis element to 0. The matrix of Τ with respect to this basis is
(7.3.1)
1Γ,ν
"0 0
1 0
0 1
0 0'
0 0
0 0
о о
1 0
In particular, the dimension of У is k, equal to the height.
Example: Ψ is the space of all (algebraic) polynomials of degree
bounded by m (so that {V}yL0 is a basis for Ψ) and Τ is the
differentiation operator:
m-\
(7.3.2)
Γ(Σ>/*ϋ = I;V_1 = Σ (7 + ΐ)^+ι^-
The vector w = x™ has height m+ 1, and {T^w}^=0 is a basis for Ψ
(so that w is a cyclic vector).
If we want the matrix of Τ to have the form (7.3.1), we have to
normalize the basis elements, say ν ■
^T,sothatvy.+1
Tv ·.
7.3.3 Shifts are the building blocks of nilpotent systems.
Theorem (Cyclic decomposition of nilpotent operators). Let (У,Т)
be α finite-dimensional nilpotent system of height k. Then
(7.3.3)
ψ:
Θ^>
7.3. Nilpotent operators
149
where У- are Ί -invariant, and each (У-, Ту) is a standard shift.
Moreover, for every h £ N, the number n(h) of summands У- of
height h in the decomposition (7.3.3) is uniquely determined.
PROOF: We use induction on к = height [(У, Τ)}.
a. If к = 1, then Τ = 0 and any decomposition У = 0 У· into one-
dimensional subspaces will do.
b. Assume that the statement is valid for systems of height less that к
and let (У, Τ) be a finite-dimensional nilpotent system of height k.
Write Цп = кег(Г) П ТУ, and let Wout С кег(Г) be a
complementary subspace, i.e., кег(Г) = Щп®WouV
The system (ТУ, Τ) is nilpotent of height к — 1 and, by the
induction hypothesis, admits a decomposition
m
ту = @yj
7=1
into standard shifts. For every j £ [l,m] denote / = height[(^·,Г)].
Let ν. £ У j be of height Ζ., so that У- = span[r, ν ·], and observe that
{TlJ~lVj} is a basis for ^n.
Let ν · be such that ν = Γν ·, and write
r. = span[r,v.] = span[^,v.].
Let Ж^и1. = 0K/ >^ be a direct sum decomposition of Wout into one-
dimensional subspaces. The claim now is:
(7.3.4) ^ = θ^·Φ0^·
Observe that (7.3.4) is the same as (7.3.3), with some summands
written as Wv
In order to prove (7.3.4) we need to verify that the set of spaces
{УЛ7=\ U{^/}j=i is independent, and their union spans У.
Independence: Assume there is a nontrivial relation £w +£w; = 0
with Uj £ У- and w( £ Щ. Let h = max height[u ·].
150
7. Structure Theorems
If h > 1, then ΣΤΗ~ιιι] = Th~l \Yuj;+Σ™ΐ) = 0 and we obtain a
nontrivial relation between the y^s. A contradiction.
If h = 1, we obtain a nontrivial relation between elements of a basis
of кег(Г). Again a contradiction.
Spanning: Denote ^ =span[{>^,^.}], /= 1,...,/, j = 1,... ,m. Then
Г^ contains every vj9 and hence Г <% = ТУ. It folows that ^D^n,
and since it contains (by its definition) Wout, we have tf/ D кег(Г).
For arbitrary ν e У, let ν e ^ be such that Γ ν = Γ v. Then v-vG
кег(Г) С ^, so that ν e <%, and <% = Г.
Finally, to prove that η (A), the number of summands У- of height h
in the decomposition, is independent of the way it is obtained, observe
that if (W,T) is an /z-shift, the dimension of TlW is h — I if h > / + 1,
and is equal to 0 if h < /. It follows that the number of nontrivial
summands remaining in Т1У is precisely dimTly — а\тТ1+хУ, and
n(h) is the number of summands for Th~] У which disappear in ТкУ,
that is,
(7.3.5) n(h) =dimTh~]y -2dimThy + dimTh+]y.
7.3.4 Definition: A cyclic decomposition of a system (У,Т) is
a direct sum decomposition of the system into irreducible cyclic sub-
spaces, that is, irreducible subspaces of the form span[r, v].
Theorem 7.3.3 allows, for systems whose minimal polynomial
have only linear prime factors, a refinement of the canonical prime-
power decomposition to a cyclic decomposition.
If minPT = Плеа(л (x ~^)m?i^ ^еп ^е canonical prime-power
decomposition, (7.1.7), gives
(7.3.6) ^=0 кег((Г-А)тА).
Aea(r)
We apply Theorem 7.3.3 to each кег((Г — А)тя), λ in σ(Γ), denote
by Ηχ = Ηλ(Τ) the sequence of the heights obtained in the cyclic
decomposition of кег((Г — А)тя), each height h repeated ηχ(h) times,
and obtain the following theorem.
7.4. The Jordan canonical form
151
Theorem. If(y,T) is a linear system whose minimal polynomial has
only linear factors, then
(7.3.7) y= 0 rxh
XeG{T),heHx
where [Ψλ Λ, Τ) is a cyclic subspace of height h ofker((T — A)).
EXERCISES FOR SECTION 7.3
ex7.3.1 In the notation of subsection 7.3.3, prove that
к
(7.3.8) dimTly= £ (h-l)n(h).
ex7.3.2 The sequences {а\тТ1У}, and {n(h)} are (each) a complete
similarity invariant for T.
ex7.3.3 Assume F = R and minPT = Φ(χ) = χ2 + 1. Prove that the map
a + bT ι—» a + bi is a (field) isomorphism of F<1) onto C.
ex7.3.4 A unipotent matrix is a matrix Λ such that A -1 is nilpotent.
Equivalent condition is "the minimal polynomial of Λ is (x— l)m for some m" or,
if the underlying field is algebraically closed, "the spectrum of Λ reduces to
{1}".
a. Prove that an upper-triangular matrix whose diagonal entries are all 1 is
unipotent.
b. Prove that if A e <Ж{п, С) is unipotent, then AJ has polynomial growth;
more precisely: there exists a constant К such that all the entries of AJ are
bounded in absolute value by Kjm~l, where m = height[Л — /].
7.4 The Jordan canonical form
7.4.1 Bases and corresponding matrices. We consider now the
case of a cyclic linear system (У,Т) whose minimal polynomial has
the form minPT = (x-X)h with λ e F.
If vis acyclic vector, i.e., Ψ = spa η [Γ, ν], then minPTv = (χ — λ)Η,
and the set {Th}hr^ is a basis for Ψ, so that Ψ is /z-dimensional.
The system (Ψ, (T — λ)) is an /z-shift, that is, nilpotent of height /z,
and v, having height /z, is a cyclic vector for the system (Ψ, (Τ — λ)) as
152
7. Structure Theorems
well. The matrix of (T — λ) in terms of the basis ν = {(Τ — λ)^ν}Ηζ\
has the form (7.3.1) and the matrix of Τ = (Τ — λ) + λ is the h χ h
matrix
0 0]
0 0
0 0
A 0
1 λ\
that has all its diagonal entries equal to A, all the entries just below the
diagonal (assuming h> 1) are equal to 1, and all other entries are 0.
If we use such bases in every summand of the cyclic
decomposition, (7.3.7), and take as a basis for Ψ the union of these bases in
successive blocks, each ordered as above, we obtain a basis of Ψ with
respect to which the matrix of Τ is the diagonal sum of matrices of the
form (7.4.1).
Theorem (Jordan canonical form). Let (V,T) be a linear system
such that the prime factors of minPT are all linear. In particular,
(У ,T) can be an arbitrary linear system over an algebraically closed
field.
Then there is a basis ν for Ψ such that the matrix ATy of Τ with
respect to ν is a diagonal sum of matrices Jx h, λ G <τ(Γ), h G Ηχ.
EXERCISES FOR SECTION 7.4
ex7.4.1 Prove that Μλ = Y,heH h is the exponent of (x — λ) in the charac-
tristic polynomial %T[x).
*7.5 The cyclic decomposition, general case
Recall that a cyclic decomposition of a system (У, Т) is a direct sum
decomposition of the system into irreducible cyclic subspaces, that is,
irreducible subspaces of the form span [7, v].
(7.4.1)
JXh
■χ
1
0
0
0
0
я
1
0
0
0
0
я
1
0
*7.5. The cyclic decomposition, general case
153
We have seen that when the underlying field is algebraically closed,
and, more generally, if all the prime factors of minPT are linear, then
the canonical prime-power decomposition can be refined to a cyclic
decomposition. We now show that this can done for every linear system.
The summands in the canonical prime-power decomposition have
the form кег(Фш(Г)) with an irreducible polynomial Ф. We show here
that such systems (whose minimal polynomial is Фш, with irreducible
Ф) admit a cyclic decomposition.
In section 7.3 we proved the claim for nilpotent operators, that is,
the special case in which Φ(χ) = χ (and which extends immediately to
the general linear case Φ(χ) = χ — A for some A E F).
The proof given below repeats, essentially verbatim, the proof
given for the nilpotent case. Observe that the nilpotent operator now
ϊϊΦ(Τ).
7.5.1 We assume that minPT = Фт and Φ is irreducible of degree d.
For every vEf, minPT v = Φ^ν), 1 < к < m, and maxv&(v) = m; we
refer to k(v) as the Φ-height, or simply height, of v, and to m as the
Φ-height, or simply height, of the system.
Theorem. There exist vectors ν Ε У such that Ψ = 0 span [Γ, ν].
Moreover, the set of the Φ-heights of the ν -'s is uniquely determined.
PROOF: We use induction on the Φ-height m.
a. m = 1. The system is semisimple and the canonical
decomposition for it is cyclic. See section 7.2.
b. Assume that minPT = Фш, m > 1, and the statement of the theorem
valid for heights lower than m.
Write Щп = кег(Ф(Т))Г)Ф(Т)У. The system (кег(Ф(Г)),Г) is
semisimple, see section 7.2, so that Щп reduces кег(Ф(Г)): there
exists a Г-invariant subspace Wout С кег(Ф(Г)) complementary to Щп
i.e., such that кег(Ф(Г)) = Щп® ^out·
(Ф(Т)У,Т) is of Φ-height m — 1 and, by the induction
hypothesis, admits a decomposition Ф(Т)У = 0}Li ^ into cyclic subspaces,
Yj = span[7,Vj]. Let Vj be such that vy. = Φ(7>;·.
154
7. Structure Theorems
Write Ψ] = span[r,vy·], and let Wout = 0K/^ be a direct sum
decomposition into cyclic subspaces. The claim now is
(7.5.1) r = 0^e0^.
To prove (7.5.1) we need to show that the spaces {У;,Щ}, where
/ = 1,..., / and j = 1,..., m, are independent, and that they span Ψ.
Independence: Assume that there is a nontrivial relation
J>y. + £ w,: = 0 with и j e Tj and w- e Щ.
Let h = max Φ-heightfw ·].
If h > 1, then L<S>(T)h-lUj = <P(T)h~] (ς^ + Σ^) = 0 and we
obtain a nontrivial relation between the 7^'s. A contradiction.
If h = 1, we obtain a nontrivial relation between elements of a basis
of кег(Ф)(Г). Again a contradiction.
Spanning: Denote % = span[{^,^}], i = 1,...,/ and j = l,...,m.
Notice first that ^ D кег(Г).
Ф(Т)У contains every vj9 and hence T$/ = ГГ. For vG^, let
ν G W be such that Γ ν = Τ v. Then v-vE кег(Г) С <% so that vG^,
and <2f = Г.
Finally, just as in the previous subsection, we denote by n{h) the
number of ν,'s of Φ-height h in the decomposition. Then dn{m) =
а\тФ(Т)т-1У and, for I = 0,... ,m-2, we have, as in (7.3.5),
(7.5.2) dn(h) = а!тФ(Т)к-1У -2а{тФ(Т)кУ + а\тФ(Т)шУ.
7.5.2 The cyclic decomposition. We now refine the canonical prime-
power decomposition (7.1.7):
Г = 0кег(Ф^(Г)),
by applying Theorem 7.5.1 to each of the summands кег(Фт-''(7)) and
obtain the following theorem.
*7.5. The cyclic decomposition, general case
155
Theorem (General cyclic decomposition). Let (У,Т) be a linear
system over a field F. Let minPT = ПФШ; be the prime-power
decomposition of its minimal polynomial. Then (У ,T) admits a cyclic
decomposition
To each к corresponds an index j = j(k) such that Yk is a direct sum-
mand <9/ker(<E>W;), so that the minimal polynomial of Ту is equal to
Φι^\ for some l (к) < m.m, andm.,,, = max l (к).
j[k) J\K) JyK-s
The polynomials Ф/(^ are called the elementary divisors of T.
Remark: We defined a cyclic decomposition as one in which the sum-
mands are irreducible. The requirement of irreducibility is satisfied
automatically if the minimal polynomial is a "prime-power", i.e., has
the form Фт with irreducible Ф. If one omits this requirement and the
minimal polynomial has several relatively prime factors, we no longer
have uniqueness of the decomposition since the direct sum of cyclic
subspaces with relatively prime minimal polynomials is itself cyclic.
EXERCISES FOR SECTION 7.5
ex7.5.1 Assume minPTv = Фт with irreducible Ф. Let и е span[r,v], and
assume Φ-heightfw] = m. Prove that spa η [Г, и] = spa η [Γ, ν].
ех7.5.2 Give an example of two operators, Τ зала S in a?(C5), such that
minPT = minPs and χτ = χ3, and yet S and Τ are not similar.
ex7.5.3 Given 3 distinct irreducible polynomials Φ;· in W[x], j = 1,2,3. Let
χ = Φ]φ3φ5? ψ(χ) = фЗфЗфЗ^ and denote
У(х,ч*) = {Т:Те^(У), πίηΡτ = ψ and χτ=χ}.
Assume that {Тк}^=1 С ^(χΨ) is such that every element in ^(χΨ) is
similar to precisely one Tk. What is ΝΊ
ex7.5.4 Assume that F is a subfield of ¥x. LetBl,B2e <M{n,F) and assume
that they are Fj-similar, i.e., B2 = C^Z^Cfor some invertible С е <JK(n,¥x).
Prove that they are F-similar.
ex7.5.5 The operatrors T,S £ а?(У) are similar if and only if they have the
same elementary divisors.
156
7. Structure Theorems
*7.6 The Jordan canonical form, general case
7.6.1 Bases and corresponding matrices. Let {Ψ, Τ) be cyclic, that
is, У = spa η [Γ, ν], and minPT = minPTv = Фш, with Φ irreducible of
degree d. The cyclic decomposition provides several natural bases:
i. The (ordered) set {Th}dr^1 is a basis; the matrix of Τ with
respect to this basis is the companion matrix of Фт.
ii. Another natural basis in this context is
(7.6.1) {r*v}£du W)rv}bdu * · ·u W)m~ T v}to;
the matrix Афт of Τ relative to this ordered basis consists of m copies
of the companion matrix of Φ arranged on the diagonal, with 1 's in the
unused positions in the sub-diagonal.
If Αφ is the companion matrix of Φ, then the matrix Αφ4, e.g., is
(7.6.2)
-φ4
Άφ
1
Άφ
1
Φ
1
Άφ
7.6.2 The canonical form for real vector spaces. When (У,Т)
is defined over M, the irreducible factors Φ of minPT are either linear
or quadratic, i.e., have the form
Φ(χ)=χ — λ, or Φ(χ) =x2 + 2bx + c with b2-c<0.
The companion matrix in the quadratic case is
(7.6.3)
0 -c
1 -2b
*7.6. The Jordan canonical form, general case
157
(OverC wehave;t2 + 2fct+c= (χ — λ)(χ — λ) with λ = —b + \/b2 — c,
and the matrix is similar to the diagonal matrix with λ and λ on the
diagonal.)
EXERCISES FOR SECTION 7.6
ex7.6.1 Prove that every square complex-valued matrix is the product DU',
where D is diagonal and U unipotent.
ex7.6.2 Assume that vx,..., vk are eigenvectors of Τ with the associated
eigenvalues λι,...,λίζ all distinct. Prove that vl,...,vk are linearly
independent.
ex7.6.3 Show that if we allow complex coefficients, the matrix (7.6.3) is
ΓΑ 01
similar to
0 A
with A
-ь+Vb2
-c.
ex7.6.4 Assume that Τ is given by the matrix AT
F3.
Γ0
1
0
0
0
1
2
0
0_
acting on
a. What is the basic decomposition when F = C, when F = R, and when
F = Q?
b. Prove that when F = Q, every nonzero vector is cyclic. Hence, every
nonzero rational vector is cyclic when F = R or C.
c. What happens to the basic decomposition under the action of an operator
S that commutes with ΤΊ
d. Describe the set of matrices A G J^(3;F) that commute with AT, where
F = C, R, and Q respectively.
ex7.6.5 Prove that the matrix
if the underlying field is R, and is diagonalizable over C. Why doesn't this
contradict exercise ex7.5.4?
-b Vc-b2]
is not similar to a triangular matrix
ex7.6.6 If b2 - с < 0, then the (real) matrices
are similar.
-Vc^F
and
-c
-2b
ex7.6.7 Let Л e Ji{n\C) such that {Aj : j e Щ is bounded (under any norm
on J%(n\C)\ in particular: all the entries are uniformly bounded). Prove that
158
7. Structure Theorems
all the eigenvalues of A are of absolute value not greater than 1. Moreover,
if Я £ σ(Λ) and |A| = 1, there are no ones under λ in the Jordan canonical
form of Λ.
ex7.6.8 Let A e Jt(n\C) such that {Aj : j e Z} is bounded. Prove that A is
diagonalizable, and all its eigenvalues have absolute value 1.
ex7.6.9 Show that, with A e Jt{n\C), the condition that {Aj : j e N} is
bounded is not sufficient to guarantee that A is diagonalizable.
However, if for some constant С and all polynomials Ρ G C[z], we have
||^(A)|| < Csup, ^^(z)!, then A is diagonalizable and all its eigenvalues
have absolute values < 1.
ex7.6.10 Let Τ e 3?(У). Write χτ = ЦФт^ where Φ;· are irreducible, but
not necessarily distinct, and m · are the corresponding heights in the cyclic
decomposition of the system.
Find a basis of the form (7.6.1) for each of the components and describe
the matrix of Τ relative to this basis.
ex7.6.11 Let A = ]χ h be the h χ h matrix defined in (7.4.1). Compute An for
all η > 1.
Hint: Write A = A/+ β.
Chapter 8
Additional Topics
Unless stated explicitly otherwise, the underlying field F of the
vector spaces discussed in this chapter is either R or C.
8.1 Functions of an operator
We assume in this section that the underlying field is C.
If Ρ = Y^CLjXJ is a polynomial with coefficients in F, we defined
P(T) by
P(T) = ZajTJ.
The map Ρ н-> Р(Т) is a homomorphism of F[jc] onto a subalgebra of
Jf(y). We can often extend the homomorphism to a larger function
space, but in most cases the range stays the same. The advantage will
be in having a better match with the natural notation arising in
applications.
8.1.1 Spectral approach. Write minPT(z) = 1\λβσ{τ)(ζ - А)ш(Я)
and observe that a necessary and sufficient condition for a polynomial
Q to be divisible by minPT is that Q be divisible by (z — A)WW for
every A G <τ(Γ), that is, have a zero of order at least m(A) at A. It
follows that P{ (T) = P2(T) if and only if the Taylor expansion of the
two polynomials are the same up to, and including, the term of order
m(A) - 1 at every λ G σ(Τ).
In particular, if m(A) = 1 for all A G σ(Τ) (i.e., if (У, T) is semisim-
ple), the condition Ρχ(λ) = Ρ2(λ) for all A G o{T) is equivalent to
Pl(T) = P2(T).
If F is an arbitrary numerical function defined on σ(Γ) and Τ is
semisimple, then the only consistent way to define F(T) is by setting
159
160
8. Additional Topics
F{T) = P(T), where Ρ is any polynomial that takes the same values as
F at each point of σ(Γ). This defines a homomorphism of the space
of all numerical functions on σ(Γ) onto the (same old) subalgebra
generated by Τ in 3?(У).
In the general case, F needs to be defined and differentiable at
least m(A) — 1 times in a neighborhood of every λ G o(T). With this
assumption, we define F(T) = P(T), where Ρ is a polynomial whose
Taylor expansion is the same as that of F up to, and including, the term
of order m(A) — 1 at every A G o(T).
8.1.2 Analytic approach. The discussion in the previous
subsection can only be put to use in practice if one has the complete spectral
information about T. One needs to have the zeros of its minimal
polynomial, including their multiplicities, given explicitly.
One can often define F(T) without explicit knowledge of this
information if F holomorphic in a sufficiently large set, and always if
F is an entire function, that is, a function that admits a power
series representation in the entire complex plane. This is done formally
just as it was for polynomials, namely, for F(z) = Σ1οαηΖη, we write
F(T) = Y^anTn, which is well defined (see* 1.5 and*2.6) if the series
converges. Since Л£(У) is finite-dimensional, so that all the norms
on it are equivalent, we can use a submultiplicative "operator norm"
as defined by (2.6.1). This keeps the estimates a little cleaner since
11 Tn 11 < || Τ ||", and if the radius of convergence of the series is larger
than ||Γ||, the convergence of Y^anTn is assured.
Examples:
a. Assume that the norm used is submultiplicative, and \\T\\ < 1; then
(/ - T) is invertible and (/- T)~l = Σ7=ο Τη.
b. Define ezT = Σ^Τη. The series is clearly convergent for every
Τ G ££(Ψ) and a G C. As a function of the parameter ζ it has the
usual properties of the exponential function.
One may be tempted to ask whether ezT has the same property as
a function of 7, that is, if ez{?+s"> = ezTezS.
8.1. Functions of an operator
161
The answer is yes if S and Τ commute, but no in general; see
exercise ex8.1.4.
8.1.3 In the context of the previous subsection, let SN(T) = Σ$ αηΤη.
Division with remainder of SN by minPT gives
(8.1.1) SN = qNm\nPT+PN,
with degPN < m = degminPT, and we have SN(T) = PN(T).
As \\PN(T)-PM(T)\\ = \\SN(T) -SM(T)\\, the sequence {PN(T)}
is a Cauchy sequence, and hence it converges to a limit P(T) G Jzf (^).
As £?{Τ) is a closed subspace of (the finite-dimensional) Jzf (^), we
have P(T) G &(T), i.e., Ρ(Γ) is a polynomial in Г of degree bounded
by m — 1, and
(8.1.2) F(7) = lim SN(T) = lim PN(T) = P(T).
The proof of the following proposition is left as an exercise.
Proposition. Let F and Τ be as above and let Ρ be the polynomial
defined by (8.1.2). Then for all λ G σ(Τ)
(8.1.3) pW(A) = F«(A), forO<k<m(X)-l.
The proposition shows that Ρ is the same polynomial that we
obtain using the spectral approach. Thus, the analytic approach can be
thought of as a practical method for computing the value of F(T) for
analytic F when the spectral information of Τ is incomplete.
EXERCISES FOR SECTION 8.1
Assume that Ψ is a finite-dimensional complex vector space.
ex8.1.1 Prove Proposition 8.1.3.
ex8.1.2 An operator Τ G if (У) has a square root if there is S G if (У) such
thatr = S2.
a. Prove that every semisimple operator on Cn has a square root.
b. Prove that every invertible operator on a finite-dimensional complex
vector space has a square root.
162
8. Additional Topics
c. Prove that the standard shift on Cn does not have a square root.
d. Let Τ be the standard shift on C3. Find a square root for / + T.
e. How many (distinct) semisimple square roots are there for the identity
operator on an «-dimensional space У ? Can you find an operator Τ ' Ε j£f (У)
that has more square roots than the identity?
ex8.1.3 For a nonsingular Τ £ j£f (У) extend the definition of Ta from a £
Ζ to a e R in a way that guarantees that for a,beR, Ta+b = TaTb (i.e.,
guarantees that {Ta}aeR is a one parameter subgroup).
ex8.1.4 Assume T,S e а?(У). Prove that
a^ eaT ebT = e(a+b)T ^
b. Prove that if S and Τ commute, then e(r+iS) = eTes.
с Verify that e^+s) φ eTes for S =
0 0
1 0
andr =
0 1
0 0
ex8.1.5 Assume that A G JK{n\C). Show that each column v(i) of etA,
thought of as a vector of differentiable functions, is a solution of the linear
system of differential equations v'(t) = Av(t).
ex8.1.6 Let Τ denote the standard shift on C". Find log(/ + Γ).
ex8.1.7 Denote ||Г||оо = тахЯб ,rJA| (the spectral norm ofT). Prove
(8.1.4)
< liminf
n—+°o
Hint: If |A| > \\Tkp for some к e N, then the series Σο λ~ηΤη converges.
Remark: The liminf appearing in (8.1.4) is in fact a limit. To see
this, notice that an = log||rn|| is subadditive: an+m < an + am. This
implies akn < kan, or -^akn < \an, for all к G N. This, in turn, implies
Y\m±an
liminf ±an.
8.2 Quadratic forms
8.2.1 A quadratic form in η variables is a polynomial Q€F[xl,...,xn]
of the form
(8.2.1) Q(xx,... ,*„) = ΣαιΛχίχ1'
Since xtXj = XjXj, there is no loss of generality in assuming a( = a- v
8.2. Quadratic forms
163
A Hermitian form on an η-dimensional inner-product space Ж is a
function of the form Q[y) = (Γν, v) with Τ G af(jff).
A basis ν = {vj,..., v„} transforms Q into a function Qy of η
variables on the underlying field, R or С as the case may be. We use the
notation appropriate1 for C.
Write ν = YUxjVj and atj = (Γν-,ν,·); then (Γν,ν) = Ztj^j^j
and
(8.2.2)
Qy(xv...,xn) = y£iaijXiXj
hj
expresses Q in terms of the variables {x·} (i.e., the v-coordinates of v).
Conversely, given a function of the form (8.2.2), denote the matrix
of coefficients (a{ ) by Av, write χ =
, and observe that
(8.2.3)
QY(xv...,xn) = (Αχ,χ) =xI?Avx.
8.2.2 If we replace the basis ν by another, say w, the coefficients
undergo a linear change of variables: there exists a matrix С G Ж(п)
У\
[ynj
of a
that transforms by left-multiplication the w-coordinates у =
vector into its v-coordinates: χ = Су. Now
(8.2.4) βν(*!,... ,xn) = S*Avx = f С*АУС у
and the matrix representing Q in terms of the variables у ·, is'
(8.2.5) Aw = C*AyC = C*AyC.
Definition: The matrices A and В are congruent if there exists a
non-singular matrix С such that В = С*AC.
Λ/of/ce fftaf the form now is C*AC, rather than C~lAC (which defines
similarity). The two notions agree if С is unitary, since then C* =C~l.
If the underlying field is Ш, the complex conjugation can simply be ignored.
The adjoint of a matrix is introduced in 6.2.3.
164
8. Additional Topics
8.2.3 Real-valued quadratic forms. When the underlying field
is R, the quadratic form Q is real-valued. It does not determine the
entries α· · uniquely. Since *·*· = *·*·, the value of Q depends on
a{ : + α · ι and not on each of the summands separately. We may
therefore assume, without modifying <2, that a{ = a- ·, thereby making the
matrix Av = (α· ·) symmetric.
For real-valued quadratic forms over С the following lemma
guarantees that the matrix of coefficients is Hermitian.
Lemma. A quadratic form x^AyX on Cn is real-valued if and only if
the matrix of coefficients Av is Hermitian,3 i.e., a{ = a~.
PROOF: If я· · = a~ for all ij, then £· ·α· jxpT- is it own complex
conjugate.
Conversely, if we assume that £· <яг ·*·*τ G Μ for all ^,... ,xn G C,
then:
Taking χ = 0 for j φ к, and xk = 1, we obtain ahgR. Taking
xk= xl = I and * · = 0 for 7 φ &, /, we obtain akl + alk eR, that is,
3^ z = —3αζ Λ; while for ^ = i, ^ = 1 we obtain i{ak i—dlh) GK,
i.e., 91^ l = 9ΐβ/ r Combining the two we have ak l = afj^. <
8.2.4 The fact that the matrix of coefficients of a real-valued quadratic
form Q is self-adjoint makes it possible to simplify Q by a (unitary)
change of variables that reduces it to a linear combination of squares.
If the given matrix is A, we invoke the spectral theorem, Theorem
6.5.3, to obtain a unitary matrix U such that U*AU = U~lAU is a
diagonal matrix whose diagonal consists of the complete collection,
including multiplicity, of the eigenvalues {A·} of A. In other words, if
χ = f/y, then
(8.2.6) Q(xv...,xn) = £Xj\yj\2.
There are other matrices С which diagonalize <2, and the
coefficients in the diagonal representation Q(yx,... ,yn) = Σ^/b/l2 depend
on the one used. What does not depend on the particular choice of С
3Equivalently, if the operator Τ is self-adjoint.
8.2. Quadratic forms
165
is the number n+ of positive coefficients, the number n0 of zeros and
the number n_ of negative coefficients. This is known as The law of
inertia.
Definition: A quadratic form Q{y) on a (real or complex) vector
space Ψ is positive definite if Q{v) > 0 for all ν φ 0 in Ψ\ it is negative
definite if Q{y) < 0 for all ν φ 0 in Ψ.
If Ψ is an inner-product space and Q{y) = (Αν,ν) with a self-
adjoint operator A, our current definition is consistent with the
definition in 6.6.1: the operator A is positive if Q(v) = (Αν,ν) is positive
definite. We use the term positive definite to avoid confusion with
positive matrices as defined in the following section.
Denote by n+ the maximal dimension of subspaces Ψχ of Ψ on
which Q is positive definite, by n_ the maximal dimension of sub-
spaces Ψχ of Ψ on which Q is negative definite.
Proposition. Let у be α basis in terms of which
Q(yv-,yn)=Lbj\yj\2>
and arrange the coordinates so that b > Ofor j <m and b < Ofor
j > m. Then m = n+.
PROOF: Denote Ψ+ = span[vp...vm], and У<0 = span[vm+1,...v„],
the complementary subspace.
0,(У\->--->Уп) is clearly positive on Ψ+, so that m < n+. On the
other hand, by Theorem 2.5.3, every subspace W of dimension > m
has elements ν G У<0, and for such ν we clearly have Q(v) < 0. <
The proposition applied to —Q shows that n_ equals the number
of negative b 's. This proves
Theorem (Law of inertia). Let Q be a real-valued quadratic form.
Then in any representation Q(yv...,yn) = ££,|;у,|2, the number of
positive coefficients is n+, the number of negative coefficients is n_,
and the number of zeros is n0 = n — n+ — n_.
166
8. Additional Topics
EXERCISES FOR SECTION 8.2
ex8.2.1 Prove that if (Av, v) = (5v, v) for all ν G Rn, with А, В G ^(/i,K),
and both symmetric, then A = B.
ex8.2.2 Let {v;·} С Ж. Write α·;· = (ν·, ν ■) and A = (αί;). Prove that A is
positive definite if and only if {v } is linearly independent.
ex8.2.3 The Gram determinant |Γ| of the vectors ν , j = 1,..., m in Ж is the
detrminant of the matrix
r = r(v1,...,vm) =
(Wi) (vi,v2) (v1?vm)'
(V2>V1> (V2'V2> (V2>V™)
.(Vm^) (vm,V2) (vm,Vm)J
a. Identify the inequality Γ(ν{, ν2) > 0.
b. Prove that |Γ| > 0, and it vanishes if and only if the vectors ν are linearly
dependent.
Hint: Identify the quadratic form (Γχ,χ).
8.3 Perron-Frobenius theory
Matrices with positive, or more generally, nonnegative coefficients
have wide application in a variety of fields. The basic spectral
properties for such matrices are described by the theorems of Perron and
Frobenius.
8.3.1 Notation and terminology.
Definition: A matrix A e Jt{n, C) is positive if all its entries are
positive.4 A is nonnegative if all its entries are nonnegative.
Similarly, a vector ν G Cn is positive if all its entries are positive,
and nonnegative, if all its entries are nonnegative.
With A denoting either matrices or vectors, Ax > A2, Αχ ^ A2, and
A{> A2 will mean respectively that A{ — A2 is nonnegative,
nonnegative but not zero, and positive.
Observe that if A > 0 and ν > 0, then Av > 0.
4Not to be confused with positivity, as defined in 6.6.1, of the operator TA of
multiplication by A.
8.3. Perron-Frobenius theory
167
The spectral norm of a matrix A G Jt{m, C) is defined by
||А||5р = тах{|т|:теа(А)}.
In this section, it will be useful (and simplify notation) to think of
Cn as the algebra of functions on the finite set [1,... ,n], and we write
vectors as ν = (v( 1),..., v{n)). The support of a nonnegative vector ν is
the set of indices j G [1,..., n] with v(j) > 0. For ν G Cn, we denote by
|v| the vector (|v(l)|,..., |v(n)|). If |v| > 0 we denote by argvande'argv
the vectors (argv(l),...,argv(w)) and (e'^1),... Уаг^М),
respectively. Also, for w, ν G Cn we denote the vector (m(1)v(1),.. .,u(n)v(n))
by mv, (this should not be confused with the inner product).
8.3.2 Positive matrices. For a positive matrix A, we denote by ρ (A)
the set of all positive numbers μ for which there exist vectors ν ^ 0
such that
(8.3.1) Αν>μν.
It is not hard to see that, on the one hand, min^a· г е р(А), and on the
other, every μ G ρ (A) is bounded by Y<ijai ,· Hence ρ (A) is nonempty
and bounded. Let ρ = sup ,дч μ.
Lemma, ρ is an eigenvalue of A and has a positive corresponding
eigenvector
PROOF: Let μη G ρ (A) be such that μη —> ρ, and let vn ^ 0 be such
that Avn > μηνη. We write, vn = (v„(l),..., v„(m)), and normalize vn
by the condition Y<jvn(j) = 1. Since now 0 < vn(j) < 1 for all η and
7, we can choose a sequence nk such that vn (j) converges for each
I < j <m. Denote the limits by v*(j) and let v* = (v*(1),...,у*(т)).
We have JV.v*(j) = 1, and since all the entries of AvHk converge to the
corresponding entries in Av*, we also have
(8.3.2) Av*>pv*,
sop ep(A).
If one of the entries in pv*, say pv*(l), were smaller than the Z'th
entry in Av*, we could replace v* by v** = v* + eel (where el is the unit
168
8. Additional Topics
vector that has 1 as its Z'th entry and zero everywhere else) with ε > 0
small enough to have
Av*(/) >pv**(/).
Since Ael is (strictly) positive, we would have Av** > Av* > pv**, and
for δ > 0 sufficiently small we would have
Αν** > (p + <5)v**
contradicting the definition of p.
This shows that the inequality in (8.3.2) is in fact an equality, so ρ
is an eigenvalue and v* is a corresponding eigenvector. Since v* ^ 0
and A > 0, it follows that pv* = Av* > 0, so v* > 0. <
8.3.3 Definition: An eigenvalue ρ of a matrix A is called
dominant if
a. p is simple: ker((A — p)2) = ker(A — p) and dimker(A — p) = 1;
and
b. every other eigenvalue Я of A satisfies |A| < |p|.
Notice that b implies that \p\ = \\A\\sp.
Theorem (Perron). Let A = (a. ) be α positive matrix. Then it has a
positive dominant eigenvalue and a positive corresponding
eigenvector Moreover, up to scalar multiplication, there is no other nonnega-
tive eigenvector for A.
PROOF: Let ρ = sup ^μ, then by Lemma 8.3.2, ρ is an
eigenvalue of A with positive eigenvector v*. Let λ φ ρ be another
eigenvalue of A and w a corresponding eigenvector. The adjoint A* = A^ is
a positive matrix and clearly, ρ is also an eigenvalue of A* with a
positive eigenvector, v*. But (w,v*) = 0 (see exercise ex6.2.4), and since
v* is strictly positive, w cannot be nonnegative. Thus, to complete the
proof of the theorem, it suffices to show that ρ is dominant.
First, we show that dim ker(A — p) = 1. If Au = pu for some vector
w, then the real and imaginary parts of и satisfy the same equality, so it
8.3. Perron-Frobenius theory
169
suffices to show that if и has real entries then it is a constant multiple
of v*. Since v* > 0, there exists a constant c^O such that v* + cu
has all its entries nonnegative, and at least one vanishing entry. Now,
v* + cu is an eigenvector for ρ and, unless v* + cu = 0, we would have
p(v* + cu) = A(v* + cu) > 0; this shows that v* + cu = 0 and и is a
multiple of v*.
Next, we show that ker((A — p)2) = ker(A — p). Assume the
contrary, and let и G ker((A — p)2) \ ker(A — p), so that (A — p)w is a
nonzero element in ker(A — p). We have
(8.3.3) Au = pu + cv*
with с ф 0. Split (8.3.3) into its real and imaginary parts:
(8.3.4) A9u* = p9u* + %:v* A3w = p3w + 3cv*.
Either Cj = %: φ 0 or c2 = Зс Φ 0 (or both). This shows that there is
no loss of generality in assuming that и and с in (8.3.3) are real-valued.
Replace u, if necessary, by ux = —u to obtain Aux = pu{ H-qv*
with сj > 0. Since v* > 0, we can choose a > 0 large enough to
guarantee that и j +m>* > 0, and observe that
A(ux +av*) = p(u{ +m>*) + c1v*
so that А (и j +av*) > p(u{ +av*) contradicting the maximality of p.
Finally, let λ be an eigenvalue of A, and let w φ 0 be a
corresponding eigenvector: Aw = Aw. Denote \w\ = (|w(l)|,..., |w(m)|).
The positivity of A implies A \w\ > \Aw\ and
(8.3.5) A\w\ > \Aw\ > \X\\w\
so that |A| G p(A), i.e., |A| < p. If |A| = ρ we must have equality in
(8.3.5) and \w\ = cv*. Equality in (8.3.5) can only happen if A\w\ =
\Aw\, which means that all the entries in w have the same argument,
i.e., w = e"*|w|. In other words, wis a constant multiple of v*, so
A=p. <
170
8. Additional Topics
8.3.4 Nonnegative matrices. Let Ш denote the matrix all of whose
entries are 1. If A > 0 then A + ^Ш> О and has, by Perron's theorem, a
dominant eigenvalue pw and a corresponding positive eigenvector vw,
which we normalize by the condition Σ!)=ι vm(j) = 1 ·
pw is monotone nonincreasing asm->oo and converges to a limit
ρ = \\A\\sp > 0 (see *A.6.9). For a sequence {m·} the vectors vm
converge to a nonnegative vector v* which, by continuity, is an eigenvector
for ρ and satisfies £v*(Z) = 1.
That is, if A is a nonnegative matrix, then \\A\\sp is an eigenvalue
of A with nonnegative eigenvector v*. However, in general, there is no
guarantee that \\A\\sp is dominant, nor indeed positive. Consider the
following examples:
a. The identity matrix. 1 is the only eigenvalue and its multiplicity is
equal to n.
b. The nilpotent matrix having ones below the diagonal, and zeros
elsewhere. The spectrum is {0}.
с The matrix Ασ of a permutation σ G Sn. The spectrum depends
on the decomposition of σ into cycles. If σ is a single cycle (of
full length), then the spectrum of Ασ is the set of roots of unity of
order n. The eigenvalue 1 has (1,..., 1) as a unique eigenvector. If
the decomposition of σ consists of к cycles (including the trivial
cycles) of lengths / > 1, j = 1,... ,&, then the spectrum of Ασ is
the union of the sets of roots of unity of order / . The eigenvalue 1
now has multiplicity /c.
Thus, for a general nonnegative matrix A
1. \\A\\sp may be zero;
2. \\A\\sp may have high multiplicity;
3. \\A\\sp may not have positive eigenvectors;
4. There may be other eigenvalues of modulus ||A||J/?.
For a transitive nonnegative matrix A, the first three problems
disappear, and the set of eigenvalues with modulus \\A\\sp has a simple
8.3. Perron-Frobenius theory
171
algebraic structure. This is the subject of Frobenius' theorem, which
follows below.
8.3.5 Definitions. Assume A > 0. We use the following
terminology:
A connects the index j to / (connects (7, /)) directlyif a{ · φ 0. Since
Ae- = Σαΐ jei> A connects (7,/) directly if e· appears (with a nonzero
coefficient) in the expansion of Ae .
More generally, A connects j to / (connects (7, /)) if, for some
positive integer k, Ak connects 7 to / directly. This means: there is a
connecting chain for (7, i), that is, a sequence {s/}f=0 suc^ ^а* J = so>
i = sk and Π?=ι β^/_ι 7^ 0. Notice that if a connecting chain for (7, /),
/ Φ 7, has two occurrences of an index &, the part of the chain
between the two is a loop that can be removed along with one к leaving
a proper chain connecting (7,/). A chain with no loops has distinct
entries and hence its length is < n. A chain which is itself a loop, that
is, connecting an index to itself, can be similarly reduced to a chain of
length < η + 1. It follows from this that A connects (7, i) if and only if
Β = ΣΖ=\ Ак connects (7,/) directly.
An index 7 is Α-recurrent if A connects it to itself—there is a
connecting chain for (7,7). The lengths к of connecting chains for (7,7)
are called return times for 7. Since connecting chains for (7,7) can be
concatenated, the set of return times for a recurrent index is an additive
semigroup of N, (a subset that is closed under addition).
The existence of a recurrent index guarantees that Aw φ 0 for all
m; in other words—A is not nilpotent. This eliminates possibility 1
above.
8.3.6 Transitive matrices. The matrix A is transitive (also called
ergodic or irreducible) if it connects every pair (7, /). If A is a nonneg-
ative transitive matrix, every index is Α-recurrent, A is not nilpotent,
andp = ||A||jp>0.
Since A connects (7,/) if and only if Β = Σ%=\ Ак connects (7,/)
directly, it follows that a nonnegative matrix A is transitive if and only
if В is positive. Since, by 8.3.4, ρ is an eigenvalue for A, it follows that
β = Σ!\Ρ^ is an eigenvalue for B, having the same eigenvector v*.
172
8. Additional Topics
Applying Perron's theorem to B, it follows that v* is in fact
positive, every other nonnegative eigenvector of В is a multiple of v* and
β is the dominant eigenvalue for B. We note that β = \\B\\sp also
follows directly from the spectral mapping theorem and the fact that
P = \\A\\sP.
Proposition. If A is a transitive nonnegative matrix, then ρ = \\A\\sp
is a simple eigenvalue of A and has a positive eigenvector v*, which is
the only nonnegative eigenvector of A, up to scalar multiplication.
PROOF: Every eigenvector of A is an eigenvector of B, so any non-
negative eigenvector of A must be a multiple of v*, and it follows that
ker(A —p) = ker(B — β) is 1-dimensional.
Next, observe that Β-β= Σ)=\ Ay -Pj = (A -p)Q(A), where Q
is a polynomial of degree η — 1. This implies that
ker((A-p)2) С ker((£-/3)2) = ker((fi-j8)) = ker((A-p)).
But ker((A - p)) С ker((A - p)2), so ker((A - p)2) = ker((A - p)).
8.3.7 Let A be a transitive nonnegative matrix, and let ρ = \\A\\sp.
By the proposition above, ρ is a simple eigenvalue of A, but it need
not be dominant: A may have other eigenvalues of the same modulus.
Without loss of generality,5 we may assume that ρ = 1, and
normalize the corresponding positive eigenvector v* = (v*(l),..., v*(n))
of A by the condition Ey v*(y) = 1.
Lemma. Assume that A is transitive, ν > 0, μ > 0, Αν ^ μ v. Then
there exists a positive vector u>v such that Au > μ^.
Proof: As in the proof of Perron's theorem: let / be an index such
that Av(l) >μν(Ζ), letO< εχ <Αν(/)-μν(Ζ) and Vj = ν + ε^.
Then Αν > μν1? hence
Avj =Av + £1Ae/ > μν! +e}Aen
By considering the transitive nonnegative matrix A = ρ ιΑ.
8.3. Perron-Frobenius theory
173
and Avj is strictly bigger than μν1 at / and at all the entries on which
Ael is positive, that is, the Vs such that a{ l > 0.
Now take v2 = vx + s2Aet with ε2 > 0 sufficiently small so that
Avj > μν2 and observe that
Av2 = Avj + e2A2ej > μν2 + e2A2ez
and Av2 — μν2 is positive on all the entries on which Avx — μν1 is
positive, as well as on the support of A2ez; in particular on / and the
support of Ael + A2ez. Continue in the same manner, taking ε3 > 0 small
enough, and v3 = v2 + £3(Ael + A2ez), so that Av2 > μν3 with strict
inequality on the support of (/ +A +A2 +A3)e/9 etc. The transitivity
of A guarantees that after к < η such steps we obtain и = vk > 0 such
that Am > μu. ^
The lemma implies in particular that if, for some μ > 0, there
exists a vector ν > 0 such that Αν ^ μν, then μ < p. This is true since
the condition Au> μ и implies that (A + ^Ш)м > (1+α)μΜίθΓα>0
sufficiently small,6 and all m. This, in turn, implies that pw > (1 +ά)μ
for all m, and hence ρ > (1 + α)μ.
Proposition. Assume \\A\\sp = 1. //"ξ = ег(р /s шг eigenvalue of A and
Up is a normalized eigenvector (that is, Σ,|Μ£ С/)I = 1) corresponding
to ξ, then
a. |mJ = v*.
b. \АиЛ = A|iie|.
PROOF: A\u^\ > \Αΐ4ξ\ = |ξιιξ| = |ιιξ|.
If AI и J t^ | и J the lemma above would imply \\A\\sp > 1, contradicting
the assumption that \\A\\sp = 1. Proposition 8.3.6 now implies part a,
which in turn implies that both sides of part b are equal to ν*, and
hence are equal. <
6See 8.3.4 for the notation.
174
8. Additional Topics
8.3.8 Part b of Proposition 8.3.7 means that each entry in Au, is a
linear combination of entries of u* that have the same argument More
precisley, we have the following structure.
The set [1,..., n] is partitioned into the level sets /^, .ч such that for
every lelm:
a. argwe (/) = #(;), and
b. A maps el into span[{e^}^G/ ], where u(s) = u(s) + φ (φ =
arg£).
In particular, A maps span[{e/}/G/ ] into span^e^}^ ].
Let η = elXf/ be another eigenvalue of A, with corresponding
eigenvector иц = e'argMnv*, and let J ,k, be the level sets of [l,...,n] on
which argi^ = j(k).
A maps every e/9 for I G J^ky into span[{ew}wG7 ] where γ(ή =
у(к) + ул It follows that if I G /^, л Π/^, then
Ael G span[{ej^w] nspan[{ew}mG</^],
where u(s) = u(j) + φ and γ(ή = y(k) + ул If we denote by u.
the vector /arg^+arg^4, then
argAe'^^+^e, = argM, + arg^ + φ + ул
This being true for all / G [1,..., n], it follows that Am, = ξ η м, .
Thus, if ξ and η are arbitrary eigenvalues of modulus 1 of A, then
ξη = e'(9+v) is aiso an eigenvalue of A. This means that the set
σ(Α)* = σ(Α)Π{ζ:|ζ| = 1}
is a finite subgroup of the multiplicative group T* of complex numbers
of modulus 1. As such, σ(Α)* is the group of roots of unity of order
m, where m is equal to the cardinality of σ(Α)*.
For general transitive nonnegative A (i.e., when ρ Φ 1) we write
σ(Α)* = {У : peu G σ(Α)}. In either case, σ(Α)* is called the period
group of A, and its order m is the periodicity of A.
If ξ = e27n/w is the generator of the period group of A and u, is
an eigenvector corresponding to the eigenvalue ре2ш/т, we call the
8.3. Perron-Frobenius theory
175
partition of [ 1,..., n] in to the level sets /^,., of arg и μ the basic partition.
The subspaces Ψ- = span[{e/ : / G 1$, Л] are Aw-invariant and are
mapped outside of themselves by Ak unless A: is a multiple of m. The
transitivity of A now implies that the restriction of Aw to ψ. is transitive
on Ψ-, with the dominant eigenvalue 1, and v^ - = £/G/ v* (Z)e/ is the
corresponding eigenvector.
The restriction of Aw to Ψ- has (counting multiplicities) |/^, .J — 1
eigenvalues of modulus smaller than 1. Summing for 1 < j < m and
invoking the spectral mapping theorem, Theorem 5.2.3, we see that A
has n — m eigenvalues of modulus < 1. This proves that the eigenvalues
in the period group are simple and have no generalized eigenvectors.
Combining these observations with Proposition 8.3.6 yields the
following theorem.
Theorem (Frobenius). Let A be α transitive nonnegative η χ η
matrix. Then ρ = \\A\\sp is a simple eigenvalue of A and has a positive
eigenvector v*, which up to scalar multiplication is the only
nonnegative eigenvector of A. Furthermore, the eigenvalues of modulus ρ are
all simple, and the set σ(Α)* = {elt : pelt G о (A)} is the group of roots
of unity of order m = |σ(Α)*|.
8.3.9 Definition: A matrix A > 0 is strongly transitive if Aw is
transitive for all m G [1,..., n].
Theorem. If A is strongly transitive, then \\A\\sp is a dominant
eigenvalue for A, and has a positive corresponding eigenvector.
PROOF: The periodicity of A has to be 1. <
8.3.10 The general nonnegative case. Let A G Ж{п) be
nonnegative. We write i <A j if A connects (ij). This defines a partial order
and induces an equivalence relation in the set of Α-recurrent indices.
(The nonrecurrent indices are not equivalent to themselves, nor to
anything else.)
176
8. Additional Topics
We can reorder the indices in a way that gives each equivalence
class a consecutive bloc, and is compatible with the partial order, i.e.,
such that for nonequivalent indices, / <A j implies / < j. This ordering
is not unique: equivalent indices can be ordered arbitrarily within their
equivalence class; pairs of equivalence classes may be <A comparable
or not comparable, in which case each may precede the other;
nonrecurrent indices may be placed consistently in more than one place.
Yet, such an order gives the matrix A a "quasi-super-triangular form":
if we denote the coefficients of the "reorganized" A again by a{ , then
a{ j = 0 for / greater than the end of the bloc containing j. That means
that now A has square transitive matrices centered on the diagonal—
the squares Jl χ Jl corresponding to the equivalence classes, while the
entries on the rest of the diagonal, at the nonrecurrent indices, as well
as in the rest of the subdiagonal, are all zeros.
This reduces much of the study of the general nonnegative A to
that of transitive matrices.
EXERCISES FOR SECTION 8.3
ex8.3.1 What part of the conclusion of Perron's theorem remains valid if the
assumption is replaced by "A is similar to a positive matrix" ?
ex8.3.2 Assume Ax ^ A2 > 0, and let ρ be the dominant eigenvalues of A,.
Prove px > p2.
ex8.3.3 Let A £ Ж(п^Ж) be such that Ρ (A) > 0 for some polynomial Ρ in
R[x]. Prove that A has an eigenvalue A £ R with positive eigenvector.
Hint: Use the spectral mapping theorem.
ex8.3.4 A nonnegative matrix A is nilpotent if and only if no index is A-
recurrent.
ex8.3.5 Let A be a nonnegative matrix whose first row is positive, and let the
<A -equivalence class of ег be [1,... ,&]. Show that A has a positive eigenvalue
λ with a nonnegative eigenvector ν whose entries v(j) are positive for j in
[1,..., k] and zero for j > к, yet there may be other positive eigenvalues larger
than λ with corresponding nonnegative eigenvectors.
Γΐ 21
Hint: For the "yet" part consider A = .
8.3. Perron-Frobenius theory
177
ex8.3.6 Let Λ be a nonnegative matrix whose first row and first column are
both positive. Prove that the properties guaranteed by Perron's theorem for
positive matrices holds for Λ.
Would the same hold under the assumption that the first row and a
column other than the first are both positive?
ex8.3.7 Prove that if the elements 7^,., of the basic partition are not equal in
size, then ker(A) is nontrivial.
Hint: Show that dimker(A) > max|7^ ., \ - min|7^ .J.
ex8.3.8 Describe the transitive matrix A if the basis elements are reordered
so that the elements of the basic partition are blocs of consecutive integers in
[1,...,/!].
ex8.3.9 Prove that if A > 0 is transitive, then so is A*.
ex8.3.10 Prove that if A > 0 is transitive, ρ = \\A\\sp, and v* is the positive
eigenvector of A*, normalized by the condition (v*, v*) = 1, then for all vGCn
1 N
(8.3.6) lim -ip-Wv=(v,v>,
N—>°o/V ^
ex8.3.11 Let σ be a permutation of [1,... ,и]. Let Ασ be the η χ η matrix
whose entries a{- are defined by
(8.3.7) *,·/ = Γ lfl = 0^'
J |^0 otherwise.
What is the spectrum of Ασ, and what are the corresponding eigenvectors?
ex8.3.12 Let 1 < к < и, and let σ e Sn be the permutation consisting of the
two cycles (1,..., k) and (k+ 1,..., л), and A = Ασ as defined above. (So
that the corresponding operator on Cn maps the basis vector e{ onto £σΛ\·)
a. Describe the positive eigenvectors of A. What are the corresponding
eigenvalues?
b. Let 0 < a, b < 1. Denote by Aa b the matrix obtained from A by replacing
the /:'th and the и'т columns of A by (cik) and (cin), resp., where clk =
1 — a, ck+x k = a and all other entries (in the column) are zero; cln = b,
cfc+i n~^~b anc* a^ otner entries are zero.
Show that 1 is a simple eigenvalue of Aa b and find a positive
corresponding eigenvector. Show also that for other eigenvalues there are no nonnegative
eigenvectors.
178
8. Additional Topics
8.4 Stochastic matrices
8.4.1 A stochastic matrix is a nonnegative matrix A = [a{ .) such that
the sum of the entries in each column7 is 1:
(8.4.1) LaU = 1·
i
A probability vector is a nonnegative vector π = (p(l),..-,p(n))
such that ΣιΡ(Ι) = 1· Observe that if A is a stochastic matrix and π a
probability vector, then Απ is a probability vector.
In applications, one considers a set of possible outcomes of an
"experiment" at a given time. The outcomes are often referred to as states,
and a probability vector assigns probabilities to the various states. The
word probability is taken here in a broad sense—if one is studying the
distribution of various populations, the "probability" of a given
population is simply its proportion in the total population.
A (stationary) η-state Markov chain is a sequence {v } >0 of
probability vectors in W1, such that
(8.4.2) Vj=Avj_l=A\,
where A is an η χ η stochastic matrix.
The matrix A is the transition matrix, and the vector v0 is referred
to as the initial probability vector. The parameter j is often referred to as
time.
8.4.2 Positive transition matrix. When the transition matrix A is
positive, we get a clear description of the evolution of the Markov
chain from Perron's theorem, Theorem 8.3.3.
Condition (8.4.1) is equivalent to u*A = w*, where u* is the row
vector (1,..., 1). This means that the dominant eigenvalue for A* is 1,
hence the dominant eigenvalue for A is 1. If v* is the corresponding
(positive) eigenvector, normalized so as to be a probability vector, then
Av* = v* and hence A;v* = v* for all j.
7The action of the matrix is (left) multiplication of column vectors. The columns
of the matrix are the images of the standard basis in M.n or Cw.
8.4. Stochastic matrices
179
If w is another eigenvector (or generalized eigenvector), it is
orthogonal to w*, that is: YJ\w(j) = 0. Also, YJ\Alw(j)\ is exponentially
small (as a function of /).
If v0 is any probability vector, we write v0 = cv* + w with w in the
span of the (generalized) eigenspaces of the nondominant eigenvalues.
By the remark above, с = Lv0(y) = ^· Then Alv0 = v* +Alw and,
since Alw —> 0 as / —> <χ>, we have Λζν0 —> ν*.
Finding the vector v* amounts to solving a homogeneous system of
η equations (knowing a priori that the solution set is one-dimensional).
The observation v* = limA^, with v0 an arbitrary probability vector,
may be a fast way way to obtain a good approximation of v*.
8.4.3 Transitive transition matrix. Denote by v, the eigenvectors
of A corresponding to eigenvalues ξ of absolute value 1, normalized so
that Vj = v* is a probability vector, and |v, | = v*. If the periodicity of A
is m, then, for every probability vector v0, the sequence A;v0 is equal
to an m-periodic sequence (given by A;w0, и0 being the component
of v0 in the span of the eigenvectors corresponding to eigenvalues of
absolute value 1, all of which are m'th roots of unity) plus a sequence
that tends to zero exponentially fast.
Since every eigenvalue ξ φ 1 of absolute value 1 is an m'th root of
unity, £™ ξ1 = 0. It follows that if v0 is a probability vector, then
ι k+m
(8.4.3) - £ A\ -> v,
rnl=k+\
exponentially fast (as a function of k).
8.4.4 Reversible Markov chains. Given a nonnegative symmetric
matrix (p.■ .), we write W. = Σ/Λ/ an(^' assuming W- > 0 for all y,
ai i'. = -щ-· The matrix A = (a- ■) is stochastic since Y^a- · = 1 for all /.
We can identify the "stable distribution"—the Α-invariant vector—
by thinking in terms of "population movement". Assume that at a
given time we have population of size fe in state j and in the next unit
of time a proportion of size a{ of this population shifts to state /. The
absolute size of the population moving from j to i is a- fe so that the
180
8. Additional Topics
new distribution is given by Ab, where b is the column vector with
entries Ъ . This description applies to any stochastic matrix, and the
stable distribution is given by b which is invariant under A, Ab = b.
The easiest way to find b in the present case is to go back to the
matrix (p.■ ·) and the weights Wj. The vector w with entries W, is A-
invariant in a very strong sense. Not only is Aw = w, but the population
exchange between any two states is even:
• the population moving from / to j is: Wta-. = ρ .;
• the population moving from j to / is: W}^; = pt ■;
• the two are equal since p. . = ρ ■ ·.
8.5 Representation of finite groups
Throughout this section G will denote a finite group.
8.5.1 Definition. A representation of a group G in a vector space
У is a homomorphism, σ: g н-> Tg, of G into the group GL(>/) of
invertible elements in Jf(y). The representation is faithful if σ is in-
jective; it is finite-dimensional if the space Ψ is finite-dimensional; the
dimension of the representation is defined to be the dimension of Ψ.
A representation σ of G in Ψ makes Ψ into a G-space (У, G, σ),
that is, a vector space for which, in addition to the vector space
operations, there is an action of G on Ψ by linear maps assigned by σ. This
means that for every g G G there is an operator σ(#) = Tg G ££{Ψ),
and for all gp^GGandvG ^,
Note that the operators Tg are necessarily invertible, i.e., belong to
GL(>/), so that a G-space is simply a vector space with a given
representation of G on it. We shall use the terms G-space and representation
interchangeably.
If σ, or equivalently the action of G, is assumed known, we denote
such spaces by (У, G) and often write g instead of σ(#) or Tg.
If σ is a representation of G in Ψ, then a G-subspace of Ψ is a
vector subspace Ψ αΨ that is invariant under the action of G. This
8.5. Representation of finite groups
181
means that gw eW for all g G G and w G W. The restriction of (the
operators assigned by) σ to W is called a subrepresentation of σ.
The following discussion is valid for abelian and nonabelian groups
alike. The abelian case, however, is much simpler, and was given in
6.5.5 as an application of Theorem 6.3.6 (the spectral theorem for
commutative, self-adjoint subalgebras).
8.5.2 The dual representation. If σ is a representation of G in У,
we obtain a representation σ* of G in Ψ* by setting G*(g) = σ(#-1)*
(the adjoint of the inverse of the action of G on У). Since both g ι—> g_1
and g ι—^ g* reverse the order of factors in a product, their combination
as used above preserves the order, and we have
σ*{8χ82) = σ*(8χ)σ*(8ι)
so that σ* is in fact a homomorphism.
If the underlying field is R or C, then we will assume that the space
is equipped with an inner product. There is no loss of generality, since
an inner product may always be introduced, e.g., by declaring a given
basis to be orthonormal. We denote the G-space by Ж in these cases.
Definition: A representation σ of G in a complex inner-product
space Ж is unitary if o(g) is unitary for every g G G.
If σ is a representation in an inner-product space Ж, then the dual
representation σ* is also a representation in Ж. If σ is unitary, then
σ* = σ.
8.5.3 Let Ψχ and Ψ2 be G-spaces. We extend the actions of G on
these spaces, denoted gx and g2 respectively, to Ψχ Θ Ψ2 and Ψχ®Ψ2
by declaring
(8.5.1) g(v10v2) = g1v10g2v2 and g(vx ®v2) = gxvx ®g2v2.
We note that Jf (Ψγ,У2) = У2®У* and as such it is a G-space.
8.5.4 G-maps. Let Ψχ and Ψ2 be G-spaces. A map S: У1^1/2'\$а
G-map if it is a linear map that commutes with the action of G. This
means that for every g G G, Sgx = g25, where g. denotes the action of
182
8. Additional Topics
G on Ψ·ν In other words, a linear map 5 is a G-map if and only if the
diagram
1 7 '2
82
%
1 7 "2
is commutative.
The prefix G- can be attached to all words describing linear maps;
thus, a G-isomorphism is an isomorphism which is a G-map, etc.
If Ψχ and Ψ2 are G-spaces, then we denote by ££G{yx, У2) the space
of G-maps of Ψχ into Ψ2.
Definition: The representations (Ух, G) and (У2, G) are equivalent
if there is a G-isomorphism S : Уху-^У2, that is, if they are isomorphic
as G-spaces.
Proposition. Let S : Yx^Y2be а G-тар. Then ker(S) /5 β subrep-
resentation ofYx, and range(S) is a subrepresentation ofY2.
PROOF: If ν G ker(S), then Sgxv = g2Sv = 0, so gxv G ker(S).
Likewise, if w G range(5), then there is a v G tj such that g2w = g2Sv =
SgjV, so g2w G range(5). ^
8.5.5 Averaging, I. Let Ψ be a finite-dimensional space. For a finite
subgroup ίί С GL(>/) we write
(8.5.2) /^ = {vGf:gv = v for all gG^}.
In words: /^ is the space of all the vectors in Ψ which are invariant
under every g in if.
Theorem. The operator
(8.5.3) π9 = щ £ g
/5 β projection onto L·.
8.5. Representation of finite groups
183
PROOF: π^ is clearly the identity on /^. All we need to do is show
that range(7T^) = /^, and for that observe that if ν = -L· LgG^ gw, then
and since {gjg : g G έί} = έί, we have gj ν = v. <
8.5.6 Averaging, II. Let Ж he a finite-dimensional, complex inner-
product space, and let if be a finite subgroup of GL(Jif).
The operator β = Χ LgG^ g*g is self-adjoint and positive on Ж,
and can be used to define a new inner product:
(8.5.4) (v,u)Q = (βν,ιι) = τ-7 £ (gv,gM>,
with the corresponding norm
Since if h = {gh : g G έί } = έί, for any h G έί, we have
(8.5.5) (hv,hM>e = τ—τ £ (ghv,ghw) = (βν,ιι),
and ||hv||g = |ΜΙρ· Thus, if is a subgroup of the unitary group
corresponding to the new inner product (·, -)Q.
Denote by Жд the inner-product space obtained by replacing the
original inner product by (·, -)Q. Let {ul,..., un} be an orthonormal
basis of Ж, and {v1?...,v„} an orthonormal basis of M^q. Define
5 G GL(^) by setting Suj = v; and extending by linearity.
So defined, 5 is an isometry from Ж onto «^, with 5_1 being the
inverse isometry. Since g is unitary on Жд, for all g G έί, it follows that
S_1gS is unitary on Ж. In other words, 5 conjugates Si to a subgroup
of the unitary group U(Jif). This proves the following theorem
Theorem. Every finite subgroup ofGL(Jif) is conjugate to a
subgroup of the unitary group и(Ж).
184
8. Additional Topics
If σ : G —> GL(Jff) is a representation of a finite group in a
complex space Ж, then if = c(G) is a finite subgroup of GL(Jif). Since
we may assume that Ж is equipped with an inner product, the
previous theorem implies the following corollary.
Corollary. Every finite-dimensional representation of a finite group
in a complex vector space is equivalent to a unitary representation.
8.5.7 Let G be a finite group and Ж a finite-dimensional G-space.
Definition: A subspace ty/ с Ж is G-reducing, or reducing for
short, if it is G-invariant and has a G-invariant complement, that is,
Ж = У/ Θ Ψ, where both summands are G-invariant.
The representation (Ж, G) is irreducible if there is no nontrivial G-
invariant subspace of Ж and reducible otherwise. In the terminology
of ex5.3.13, Ж is irreducible if (Ж^) is minimal.
Proposition. Ifa:G^ СЬ(Ж) is a representation ofG in a finite-
dimensional complex vector space, then every G-invariant subspace is
reducing.
PROOF: We assume, without loss of generality, that Ж is an inner-
product space. Endow Ж with the inner product given by (8.5.4),
with if = cj(G), making the representation unitary. Now, observe that
if ty is a nontrivial G-invariant subspace, then so is its orthogonal
complement. ч
If Ж = ^0У, and both У/ and Ψ are G-invariant, then we
say that the representation {Ж\G) is the sum of the representations
(^,G)and(r,G).
Lemma. Let (У\G) and (^,G) be irreducible subrepresentations
of(Jff, G). Then, either <& Π У = {0}, orfy = У.
PROOF: У/ П Ψ is clearly G-invariant. <
Theorem. Every finite-dimensional complex representation (Ж,С)
of a finite group G is a sum of irreducible representations. That is,
(8.5.6) Ж = φ ^·,
8.5. Representation of finite groups
185
where each ^ is an irreducible G-subspace of Ж. Furthermore, the
decomposition is unique.
PROOF: We prove the existence by induction on the dimension.
If dim^f = 1, then there is nothing to prove. Assume now that
the statement is valid for representations of dimension < n, and let
(Ж, G) be a representation of dimension n.
If (Jf?,G) is irreducible, then we are done. Otherwise,
Proposition 8.5.7 shows that there is direct sum decomposition Ж = % Θ Ψ
where both summands are nontrivial G-invariant subspaces. Since ty/
and Ψ have dimensions less than n, they each have a decomposition
into irreducible G-subspaces, by the induction hypothesis. The sum of
the two decompositions is a decomposition of Ж.
The uniqueness follows from lemma 8.5.7. <
8.5.8 The regular representation. Let G be a finite group. Denote
by £2(G) the vector space of all complex-valued functions on G, and
define the inner product, for φ, ψ e £2(G),hy
(φ,ψ) = ΣφΜυΦΟ-
xeG
For gGG, the left-translation byg is the operator ρ (g) on £2(G) defined
by
Clearly p(g) is linear and, in fact, unitary. Moreover,
(Ρ(8ι82)ψ)(χ) = <Ρ((8ι82)~1χ) = <Ρ(8ι1 (ίΓ1*)) = (Ρ(«ι)Ρ(«2)φ) Μ
so that p(glg2) = P(g\)p(g2) апс* Ρ ls a unitary representation of G.
It is called the regular representation ofG.
If Η С G is a subgroup, we denote by £2(G/H) the subspace of
£2(G) of the functions that are constant on left cosets of Я.
Since multiplication on the left by arbitrary g G G maps left H-
cosets onto left Я-cosets, £2(G/H) is ρ (g)-invariant. Unless G has
no nontrivial subgroups—we say that G is simple in this case—ρ is
reducible. This proves the following proposition.
186
8. Additional Topics
Proposition. If the regular representation ofG is irreducible, then G
is simple.
The converse is false! A cyclic group of order p, with prime /?,
is simple. Yet, it follows from the discussion in subsection 6.5.5 that
the regular representation of any finite abelian group is reducible to
1-dimensional summands.
8.5.9 Let Ж be a complex G-space and let (,) be an inner product
in Ж. Fix a nonzero vector и G Ж and, for ν G Ж and g G G, define
(8.5.7) Mg) = (g~\u).
The map 5: ν н-> fv is a linear map from Ж into i2(G).
Lemma. If Ж is irreducible, then S is infective.
PROOF: If ν φ 0, then the set {gv : g G G} spans Ж. This implies that
/v is not the 0 function, i.e., 5(v) φ 0. ^
Observe that for γ£ G,
(8.5.8) p(7)/v(s) = /,(7"V) = <g-y,K> = Mg),
so that the space БЖ = ^ С ^2(G) is a G-invariant subspace, and so,
by Proposition 8.5.7, a reducing subspace of the regular representation
of f(G). Furthermore, 5 maps (Ж ,G) onto (Ж5,р\ж).
Together with the lemma above, this proves the following theorem.
Theorem. Every irreducible representation of a finite group G in a
finite-dimensional complex space is equivalent to a subrepresentation
of the regular representation.
Corollary. There are only a finite number of distinct irreducible
complex representations of a finite group G.
Appendix
A.I Equivalence relations-partitions
АЛЛ Binary relations. Formally, a binary relation R in a set Ζ is a
subset R С X x X. It is the set of all pairs x,y such that χ and у "have
relation R\ (x,y) G Λ is most often written as xRy, where R is a symbol
designating the relation.
Examples:
a. Equality: R = {(x,x) :x£X} = {(x>y) :x,y G X; * = ;y}·
& Order in Z: R = {(x,y) :x <y}.
c. Divisibility in N: R = {(m,n) :m\n},(m divides η in N).
The symbol R in each of these is the usual one, i.e., =, <, and I
respectively.
АЛ.2 Equivalence relations. An equivalence relation in a set Ζ is a
binary relation (denoted here χ = у) that is
reflexive: for all χ G Χ, χ = χ;
symmetric: for all x,y G X, if χ = у, then у = χ; and
transitive: for all x,y,z G X, if * ξ у and jez, then χ ξ ζ.
Examples:
α. Of the binary relations above, equality is an equivalence relation;
order and divisibility are not.
& Congruence modulo an integer. Here X = Z, the set of integers.
Fix an integer k. We say that χ is congruent toy modulo к and write
χ = у (mod &) if x — у is an integer multiple of k.
187
188
Appendix
с. For Ζ = {(т,и) : m,n G Ζ, η т^ 0}, define (т,и) ξ (m^Wj) by the
condition mnx = mjft. This will be familiar if we write the pairs
as ™ instead of (m,n) and observe that the condition mnl = mjft is
the one defining the equality of the rational fractions ™ and ^-.
A.1.3 Partitions. A partition of X is a collection «^ of pairwise
disjoint subsets Pa CX whose union is X, i.e.,
ΡαΠΡβ=(ΰ if α φ β, and |Jpa=Z.
A partition & defines an equivalence relation: by definition, x = y
if and only if χ and у belong to the same element of the partition.
Conversely, given an equivalence relation on X, we define the
equivalence class of χ e X as the set Sx = {y e X : χ = у}. The defining
properties of equivalence can be rephrased as:
b. If у e £x, then χ e £y, and
c. If у e £x and ζ £ <§y, then zE4
These conditions guarantee that different equivalence classes are
disjoint and the collection of all the equivalence classes is a partition
of Ζ (that defines the given equivalence relation).
EXERCISES FOR SECTION A.l
exA.1.1 Let7?! С К xR = {(χ,у) : \х-у\ < 1} andx-^ у when (;t,;y) e/^.
Is this an equivalence relation, and if not—what fails?
exA.1.2 Congruence mod к is the relation R = {(m, л) £ Ζ2 :m-ne kZ}.
Identify the equivalence classes for congruence mod k.
A.2 Maps
The terms used to describe properties of maps vary by author, by
time, by subject matter, etc. We shall use the following:
A map φ: X —> Υ is injective if jcj 7^ jc2 => ф (*!) φ ψ (^) ·
Equivalent terminology: φ is one-to-one (or 1-1), or φ is a monomorphism.
A map φ: Ζ —> Υ is surjective if <p(X) = {φ(*) : * G X} = У.
Equivalent terminology: φ is onto, or φ is an epimorphism.
A.3. Groups
189
A map φ: Χ —> Υ is bijective if it is both injective and surjective:
for every у £ У there is precisely one χ £ X such that у = φ (χ).
Bijective maps are invertible—the inverse map is defined by: φ-1 (у) = χ if
у = ф(*)·
Maps that preserve some structure are called morphisms, often
with a prefix providing additional information. Besides the mono- and
epi- mentioned above, we use systematically homomorphism,
isomorphism, etc.
A permutation of a set is a bijective map of the set onto itself.
A.3 Groups
A.3.1 Definition: A group is a pair (G, *), where G is a set and *
is a binary operation (x,y) ι—> **;y, defined for all pairs (x,y) £ GxG,
taking values in G, and satisfying the following conditions:
G-l The operation is associative: Forx^z £ G, (x*y)*z = x*(y*z).
G-2 There exists a unique element e £ G called the identity element
or the unit of G, such that e*x = x*e =x for all* £ G.
G-3 For every χ e G there exists a unique element x~l, called ffte
inverse of *, such that x~ * * * = * * x~ * = e.
A group (G,*) is abelian, or commutative, if χ ^y = у ^x for M
χ and y. The group operation in a commutative group is often written
and referred to as addition, in which case the identity element is written
as 0, and the inverse of χ as — x.
When the group operation is written as multiplication, the
operation symbol * is sometimes written as a dot (i.e., χ · у rather than χ *y)
and is often omitted altogether. We also simplify the notation by
referring to the group, when the binary operation is assumed known, as G
rather than (G, *).
The order of a group G, denoted |G|, is its cardinality.
Examples:
α. (Ζ, +), the integers with standard addition.
190
Appendix
b. (R \ {0}, ·), the non-zero real numbers, standard multiplication.
c. S„, the symmetric group on [1,..., n]. Here η is a positive integer,
the elements of S„ are all the permutations σ of the set [1,..., n],
and the operation is composition: for σ, τ Ε S„ and 1 < j < η we
set(Ta)(;) = T(a(;)).
More generally, if Ζ is a set, the collection S(X) of permutations,
i.e., invertible self-maps of X, is a group under composition. (Thus
S„ = S([1,...,/i])).
The first two examples are commutative; the third, if η > 2, is not.
A.3.2 Let G·, / = 1,2, be groups.
Definition: A map φ: Gx —> G2 is a homomorphism if
(A.3.1) ф(^) = фМф(у).
Notice that the multiplication on the left-hand side is in Gp while that
on the right-hand side is in G2.
The definition of homomorphism is quite broad; we do not
assume the mapping to be injective (1-1), nor surjective (onto). We use
the proper adjectives explicitly whenever relevant: monomorphism for
injective homomorphism and epimorphism for one that is surjective.
An isomorphism is a homomorphism which is bijective, that is, both
injective and surjective. Bijective maps are invertible, and the inverse
of an isomorphism is an isomorphism. For the proof we only have to
show that φ-1 is multiplicative (as in (A.3.1)), that is, that for g,/z Ε
G2, (p~\gh) = φ-1(#)φ-1(/ζ). But, if g = φ(χ) and h = <p(y)9 this is
equivalent to gh = (p(xy), which is the multiplicativity of φ.
If φ: Gj —> G2 and ψ: G2 ^> G3 are both isomorphisms, then
ψφ: Gx —> G3 is an isomorphism as well.
We say that two groups G and Gy are isomorphic if there is an
isomorphism of one onto the other. The discussion above makes it
clear that this is an equivalence relation.
A.3. Groups
191
A.3.3 Inner automorphisms and conjugacy classes. An
isomorphism of a group onto itself is called an automorphism. A special class
of automorphisms, the inner automorphisms, are the conjugations by
elements у G G:
(A.3.2) q>yx = y~lxy.
One checks easily (left as an exercise) that for all у G G, the map (py is
in fact an automorphism of G.
An important equivalence relation in G is conjugacy, defined by:
χ ~ ζ if there exists у G G such that ζ = (fax = у-1*?.
To check that every χ is conjugate to itself take y = e, the identity.
If ζ = «pyx, then χ = φ _,z, proving the symmetry. Finally, if ζ = y~lxy
and и = w~lzw, then
w = w~lzw = w_1^-1^vv = (;yw)_1.x(;yvv),
which proves the transitivity.
The equivalence classes defined on G by conjugation are called
conjugacy classes.
A.3.4 Subgroups and cosets.
Definition: A subgroup of a group G is a subset Η с G such that
SG-1 Я is closed under multiplication: if hl,/г2 G Я then h]h2e H.
SG-2 К^Я^Ьеп/Г1 ей.
Observe that these conditions imply that e G Я. In other words, Я С G
is a subgroup if, with the operation inherited from G, it is a group.
Examples:
a. {e}, the subset whose only term is the identity element.
b. In Z, the set qL of all the integral multiples of some integer q. This
is a special case of the following example.
192
Appendix
c. For any χ e G, the set {хк}ке% is the subgroup generated by x. A
group generated by one of its elements is called cyclic. The element
χ is of order m if the cyclic group it generates is of order m. (That is,
if m is the smallest positive integer for which xm = е.) х has infinite
order if {xn} is infinite, in which case η н-> χ? is an isomorphism
of Ζ onto the group generated by x.
d. The subset of S„ of all the permutations that leave some (fixed)
/ G [l,...,n] in its place, that is, {σ G Sn : σ(Ζ) = /}.
If φ: G ^ Gx is a homomorphism and ^ denotes the identity in Gp
then {g G G : φ# = ex} is a subgroup of G (f/?e /cerne/ of φ).
Let Я С G be a subgroup. For χ G G the set xH = {χ/ζ: h G #} is
called a /eft cosef of Я, and the set Hx = {to : h G #} is called a r/g/?f
cosef of#.
Lemma. Ifx,y G G, /7гея rte left cosets xH and yH are either
identical or disjoint. In other words, the collection of distinct xH is a
partition ofG.
PROOF: We check that the binary relation defined by "x G yH" is an
equivalence relation. The cosets xH are the elements of the
corresponding partition.
a. Reflexive: χ G xH, since x — xe and e G H.
b. Symmetric: If у G xH, then there is a z G Η such that у = xz. This
implies that χ = yz~l G yH, since Я is a subgroup.
c. Transitive: If w G yH and у G ;t#, then for appropriate hx, h2 G Я,
у = jc/Zj and w = ;у/г2 = xhxh2, and w G xH since /2^2 G Я. ^
The same proof shows that the set of right cosets of Я in G forms a
partition of G. The map хЯ н-> Hx is clearly a (set-theoretic) bijection.
The cardinality of the set of left (or right) cosets of Я in G is called
the index of Я in G, and denoted by [G: H].
A.3. Groups
193
A.3.5 Normal subgroups. Let G and Gx be groups, and φ: G^> G{
a homomorphism. Let К с G be the kernel of φ, that is, the set of
elements of G that are mapped by φ to ev the identity element of Gv
К is clearly a subgroup of G. We observe that if φ (A:) = ^ and у е G
is arbitrary, then
(p(y-lky) = (p(y)-lel(p(y) = el and /c^"1^"1)),,
so that y~lKy = A'. In other words, ^ is mapped onto itself by every
inner automorphism.
Definition: A subgroup Я с G is a normal subgroup if it is mapped
onto itself by every inner automorphism, that is, Η = y~lHy for every
Let Я с G be a subgroup and let хЯ be a left coset of #; then
xH = хЯх"1! = #jjc, which means that хЯ is a r/g/?f cosef of the group
Я1 = .хЯх-1. If Я is a normal subgroup of G, then Ηχ = Я, and it
follows that xH = Hx for all iGG.
We note this as a lemma:
Lemma. If Η С G w normal, then every left coset of Η is also a right
coset.
Proposition. The quotient G/H ofG by the equivalence relation
defined by a normal subgroup H, has a natural group structure under
which the map χ н-> xH is a homomorphism whose kernel is H.
PROOF: Let xH and yH be cosets. Define xHyH = xyH. We need to
show first that the product is well defined independently of the choice
of the representatives χ and у of the cosets. If we replace χ by another
representative xhv and у by yh2, then, by the lemma, hxy = yh3 for
some h3 G Я, and we have xh]yh2H = xyh3h2H = xyH.
The facts that this multiplication is associative, that Я is the unit
element, and that the inverse of xH isx~lH are clear, and the fact that
χ н-> xH is a homomorphism whose kernel is Я is obvious. <
194
Appendix
EXERCISES FOR SECTION A.3
exA.3.1 Check that, for any group G and every у G G, the map (pyx = y~lxy
is an automorphism of G.
exA.3.2 Let G be a finite group of order m. Let Я С G be a subgroup. Prove
that the order of Я divides m.
exA.3.3 Let Я be a normal subgroup of G and U С G a subgroup.
a. Prove that U Π Я is a normal subgroup of £/.
ft. Write £/Я = {uh:ueU,heH}. Prove that £/Я is a subgroup of G.
с Prove that UH/H is isomorphic to U/(U П Я).
*A.4 Group actions
A.4.1 Actions.
Definition: An action of G on a setX is a homomorphism φ of G
into 5(X), the group of invertible self-maps (permutations) of X.
The action defines a map (g,x) н-> <p(g)*. The notation <p(g)* is
often replaced by the simpler gx, when φ is implicitly understood.
With the simpler notation, the assumption that φ is a homomorphism
is equivalent to the conditions:
gal. ex = x for all χ G X (e is the identity element of G).
ga2. (gxg2)x = gx (g2x) for all g. eG,xeX.
Examples:
a. G acts on itself (X = G) by left multiplication: (*,;y) н^ ху.
ft. G acts on itself (X = G) by right multiplication (by the inverse):
(x,y) y->yx~l. (Remember that (ab)~l =b~la~l.)
c. G acts on itself by conjugation: (x,y) ь-> φ(χ)γ where φ(χ)γ =
xyjc-1.
d. Sn acts as mappings on [1,..., n].
*A.4. Group actions
195
A.4.2 Orbits. The orbit of an element χ e X under the action of a
group G is the set Orb (x) = {gx : g G G}.
The orbits of a G-action form a partition of X. This means that
any two orbits, Ort^Xj) and Orb(x2), are either identical (as sets) or
disjoint. In fact, if χ G Orb (y), then χ = g0y for some g0 G G, so that
y = gQlx^ndgy = gg-]x.
Since the set {gg^1 : g G G} is exactly G, we have Orb(y) =
Orb (x). If Orb (xj) Π Orb (x2) is not empty, take χ G Orb (xj HOrb (x2)
and then Orb(x) = Orb(xj) = Orb(x2). It follows that the relation
χ = у defined by: Orb (x) = Orb (y) is an equivalence relation on X.
Examples:
a. A subgroup Η С G acts on G by right multiplication: (h,g) н—> gh.
The orbit of g G G under this action is the (left) coset gH.
b. Sn acts on [1,... ,n], (cr J) ^—> o(j). Since the action is transitive,
there is a unique orbit—[1,..., n].
c. If σ G S„, the group (σ) (generated by σ) is the subgroup {σ^}
of all the powers of σ. Orbits of elements a G [1,..., n] under the
action of (σ), i.e., the sets {ak(a)}, are called cycles of σ and are
written (flj,...,^), where α·+1 = <τ(α·), and /, the period of tfj
under σ, is the first positive integer such that σι(αλ) = αλ.
Notice that cycles are "enriched orbits", that is, orbits with some
additional structure, here the cyclic order inherited from Z. This cyclic
order defines σ uniquely on the orbit, and is identified with the
permutation that agrees with σ on the elements that appear in it, and leaves
every other element in its place. For example, (1,2,5) is the
permutation that maps 1 to 2, maps 2 to 5, and 5 to 1, leaving every other
element unchanged. Notice that n, the cardinality of the complete set
on which Sn acts, does not enter the notation and is in fact irrelevant
(provided that all the entries in the cycle are bounded by it; here η > 5).
Thus, breaking [1,... ,n] into σ-orbits amounts to writing σ as a
product of disjoint cycles (see 4.1).
196
Appendix
A.4.3 Conjugation. Two actions of a group G, φχ: G χ Χχ —> Χχ,
and φ2: G χ Χ2 —> Χ2 are conjugate to each other if there is an invertible
map Ψ: Χχ —> X2 such that for all χ G G and у G Xj,
(A.4.1) φ2(χ)Ψγ = Ψ(φχ(χ)γ) or, equivalently, φ2 = Ψφ^-1.
This is often stated as: the following diagrams commute:
X,
Ψι
ψ
ψ or, equivalently, ψ-ι
Ψι
Χ,
ψ
Χο
φ2
Χο
Χ,
φ2
Χο
meaning that the composition of maps associated with arrows along
a path depends only on the starting and the end point, and not on the
path chosen.
A.5 Rings and algebras
A.5.1 Rings.
Definition: A ring is a triplet (^,+,·), where S% is a set, and +
and · are binary operations on £% called addition and multiplication
respectively, such that (^, +) is a commutative group, the
multiplication is associative (but not necessarily commutative), and the addition
and multiplication are related by the distributive laws:
a(b + c) = ab + ac and (b + c)a = ba + ca.
A subring£%x of a ring £% is a subset of £% that is a ring under the
operations induced by the ring operations, i.e., addition and multiplication,
in^.
A commutative ring with multiplicative identity is often called a
domain. A domain R in which ab = 0 implies a = 0 or b = 0 is called
an integral domain.
Examples:
a. Any field, F, is also a ring.
A.5. Rings and algebras
197
b. Ζ is an integral domain.
c. For q G Z, q > 1, the subring gZ = {gn : η G Z} is a commutative
ring without multiplicative identity.
d. The (finite) ring Z^, defined in exl.1.4, is a field if q is prime. If
q = ш, with n,m> 1, then Z^ is an example of a domain that is
not an integral domain, since η and m are nonzero and nm = 0 in
e. «y#(n) with multiplication defined by matrix multiplication, is a
ring with identity that is not commutative.
A.5.2 Ideals.
Definition: A left (resp. right) ideal in a ring £% is a subring / that is
closed under multiplication on the left (resp. right) by elements of £%\
for ae& and h G / we have ahe I (resp. ha G /). A two-sided ideal is
a subring that is both a left ideal and a right ideal.
If the ring is commutative, the adjectives "left", "right" are
irrelevant.
Assume that & has an identity element. For ge«f, the set Ig =
{ag : a G &} is a left ideal in ^, and is clearly the smallest (left) ideal
that contains g.
Ideals of the form Ig are called principal left ideals, and g is called a
generator of /^. One defines principal right ideals similarly.
A.5.3 Principal ideal domains.
DEFINITION: An integral domain in which every ideal is principal is
called a principal ideal domain, (P.I.D.).
Principal ideal domains are important in number theory in the
study of divisibility and factorization. The canonical example of a
P.I.D. is Z. The proof that Ζ is a P.I.D. follows from "division with
remainder".
Theorem (Division with remainder in Z). For every pair m,n G Z,
with m,n^0, there exists a unique pair q,r G Ζ with 0 < r < \n\, such
that
(A.5.1)
m = qn + r.
198
Appendix
PROOF: We prove the existence, and leave uniqueness as an exercise.
If m = qn, with q G Z, then the result is obvious, so we assume that
η does not divide m.
Now, assume that m,n > 0, and let r0 be the smallest positive
integer in the set S = {m — qn : q £ Z}. If r > n, then Г] = r0 — η is a
smaller nonnegative member of 5, contradicting the minimality of r0,
so 0 < r0 < n. If q0 = (m — r0)/n, then the pair g0, r0 is as required.
If m > 0 and η < 0, and g0, r0 is the pair that works for m and \n\,
then — g0 and r0 work for m and n.
If m < 0 and η > 0, and g0, r0 is the pair that works for |m| and n,
then — (q0 + 1) and η — r0 work for m and n.
Finally, if m,n < 0, and g0, r0 is the pair that works for \m\ and |n|,
then ^0 + l and η — r0 work for m and η. ^
Corollary. Ζ /s β principal ideal domain.
PROOF: Let m be the smallest positive element of / and η G /, η > 0.
By the theorem above, we can divide with remainder: η = qm + r with
g, r integers, and 0 < r < m. Since both η and qm are in /, so is r. Since
m is the smallest positive element in /, r = 0 and η = gm. Thus, all the
positive elements of / are divisible by m (and so are their negatives).
<
If m G Z, 7 = 1,2, the set/mpm^ = {η^χ +n2m2 : n1?n2 G Z} is an
ideal in Z, and hence has the form gZ. As g divides every element in
hnvm2, it divides both mx and m2; as g = nxmx +n2m2 for appropriate
η ·, every common divisor of mx and m2 divides g. It follows that g is
their greatest common divisor, g = gcd(mj ,m2). We summarize:
Proposition. Ifmx andm2 are integers, then for appropriate integers
ftp n2,
gcd(m1,m2) = nxmx +n2m2.
A.5.4 Euclidean domains. The proof that Ζ is a P.I.D. can be
repeated almost verbatim for any ring that has an appropriate notion of
division with remainder.
A.5. Rings and algebras
199
Definition: An integral domain R is called a Euclidean domain if
there exists a function ν from the set of nonzero elements in R to the
nonnegative integers that satisfies the division with remainder property
in R: if a,b G R and b φ 0, then there exist q,r G R, with v(r) < v(b)
or r = 0, such that
(A.5.2) a = qb + r.
Such a function is called a valuation on /?. We note that in general,
(i.e., in Euclidean domains other than Z), the pair q, r is not necessarily
unique.
Theorem. If R is α Euclidean domain, then R is a principal ideal
domain.
PROOF: Let / be a nontrivial ideal in R, and let a G / be an element of
minimal value in /, i.e., ν (a) < v(b) for all b φ 0 in /. Then the proof
that β is a generator of / is identical to the proof of Corollary A.5.3. <
The most important example of a Euclidean domain (besides Z) is
the ring of polynomials over a field (see A.6).
A.5.5 Fields. Fields were defined in Chapter 1. We observe that a
field may be defined equivalently as a domain in which every nonzero
element is invertible. The most basic examples are the fields of rational
numbers, Q; real numbers, IR; and complex numbers, C, which are all
infinite fields.
As mentioned in Chapter 1, there are also finite fields. In exercise
exl.1.4, we show how to construct a field with exactly ρ elements,
where ρ is a prime number. More generally, it can be shown that for
every prime ρ and every integer k>l, there is exactly one field (up to
isomorphism) with pk elements. See example d below.
Two fields, Fj and F2, are isomorphic if there is a bijective map
φ : Fj —> F2 satisfying (p(a + b) = <p(a) + <p(fe) and <p(afe) = <p(a)<p(b),
forallfljfeEFj.
Definition: If F с IK are fields, then we say that F is a subfieldof
IK, and equivalently, that IK is an extension of F. More generally, IK is
an extension of F if К contains a subfield that is isomorphic to F.
200
Appendix
An extension IK of F is also a vector space over F. The degree of
the field extension IK of F is, by definition, the dimension of IK as a
vector space over F. The extension is finite if its degree is finite.
Examples:
a. С is an extension of R of degree 2.
b. R is an extension of Q of infinite degree.
с Q(v/2) = {a + b\/2 + c\ft: a,b,c G Q} is an extension of Q of
degree 3.
d. F4 = {0, l,a,fe}, with addition and multiplication given in the
tables below, is a degree 2 extension of Z2.
+ J
~οΊ
1
a
b
L°_j
Го"
1
a
b
1
1
0
b
a
a
a
b
0
1
b
b
a
1
0
X
~οΊ
1
a
b
0
Го"
0
0
0
1
0
1
a
b
a
0
a
b
1
b
0
b
1
a
A.5.6 Algebras.
Definition: An algebra over a field F is a ring si and a
multiplication of elements of srf by scalars (elements of F), that is, a map
Fx^^j/ such that if we denote the image of [а, и) by au, we have,
for a,b G F and m,vG^,
identity: \u = u\
associativity: a(bu) = (ab)u, a(uv) = {au)v\
distributivity: {a + b)u = au + bu and а(и + ν) = au + m>.
A subalgebra six С ^ is a subring of srf that is also closed under
multiplication by scalars.
A left (resp. right, resp. two-sided) ideal in an algebra si is a
subalgebra of si that is closed under left (resp. right, resp, either left
or right) multiplication by elements of si.
A.6. Polynomials
201
Examples:
a. ¥[x], the algebra of polynomials in one variable χ with coefficients
from F, and the standard addition, multiplication, and
multiplication by scalars. It is an algebra over F.
b. С [jc, у], the (algebra of) polynomials in two variables jc, у with
complex coefficients, and the standard operations. C[x,y] is a "complex
algebra", that is, an algebra over C.
Notice that by restricting the scalar field to, say, M, a complex
algebra can be viewed as a "real algebra", i.e., an algebra over R.
The underlying field is part of the definition of an algebra. The
"complex" and the "real" C[x,y] are different algebras.
c. Ж{п), the nx η matrices with matrix multiplication as product.
EXERCISES FOR SECTION A.5
exA.5.1 Let S% be a ring with identity, and Bc^a set. Prove that the ideal
generated by B, that is, the smallest ideal / in 3% that contains B, is the set
I={L<*jbj:aje&,bjeB}.
exA.5.2 Verify that Ър is a field.
Hint: If ρ is a prime and 0 < m < ρ then gcd(m,p) = 1.
exA.5.3 Prove that the set of invertible elements in a ring with an identity is
a multiplicative group.
exA.5.4 Show that the set of polynomials {Ρ : Ρ = Σ/>2 ajxJ'} ls an ^ea* *η
¥[x], and that {Ρ : Ρ = У ·<7 α,·*·7'} is an additive subgroup but not an ideal
A.6 Polynomials
Let F be a field and F[jc] the algebra of polynomials Ρ = Σοαίχ^
in the variable χ with coefficients from F. The degree of P, deg(P),
is the highest power of χ appearing in Ρ with non-zero coefficient. If
deg(P) = n, then апх? is called the leading term of P, and an the leading
coefficient. A polynomial is called monic if its leading coefficient is 1.
202
Appendix
A.6.1 Division with remainder. By definition, an ideal in a ring is
principal if it consists of all the multiples of one of its elements, called
a generator of the ideal. The ring ¥[x] shares with Ζ the property of
being a principal ideal domain—every ideal is principal. The proof for
F[jc] is virtually the same as the one we had for Z, and is again based
on division with remainder.
Theorem. Let P,Fe F[x]. There exist polynomials Q,Re ¥[x] such
thatdeg(R) < deg(F) and
(A.6.1) P = QF + R.
Proof: Write Ρ = Σ"=οα]χ] and F = LJ=objxJ with both an and bm
nonzero, so that deg(P) = η and deg(F) = m.
If η < m, there is nothing to prove: Ρ = 0 · F + P.
If η > m, we write qn-m = an/bm, and Ρχ = Ρ — qn_mxn~mF, so
that Ρ = qn_m^~mF + Px with nx = deg^) < n.
If nx < m, we are done. If nx > m, write the leading term of Px as
aXnxn\ and set ^_w = aXnJbm, and P2 = Px -qni_m^-mF. Now
deg(P2) <deg(P1) <mdP=(qn_mx»-™ + qni_mx"i-™)F + P2.
Repeating the procedure a total of A: times, k<n — m+l,we obtain
Ρ = QF + Pk with deg(^) < m, and the statement follows with R = Pk.
<
A useful application of the theorem is the observation that a
polynomial Ρ G ¥[x] vanishes at a point λ G F if and only if Ρ is divisible
by χ — λ. Division with remainder gives Ρ = (χ — λ) Q + r, with r G F.
Evaluating at χ = Я shows that the remainder r is equal to Ρ (λ).
The same idea in a more abstract context gives the following:
Corollary. Let I С ¥[х] be an ideal, and let P0 be an element of
minimal degree in I. Then P0 is a generator for I.
PROOF: If Pel, write Ρ = QP0 + R, with deg(tf) < deg(P0). Since
R = P — QPQ g /, and 0 is the only element of/ whose degree is smaller
than deg(P0), we have Ρ = QP0. м
A.6. Polynomials
203
Since a generator of / divides every Ρ in /, it is an element of
minimal degree in /, so that a polynomial in / is a generator if and only
if it is an element of minimal degree in /.
The generator P0 is unique up to multiplication by a scalar. If
Px is another generator, each of the two divides the other, and since
they have the same degree, the quotients are scalars. It follows that
if we normalize P0 by requiring that it be monic, that is, with leading
coefficient 1, it is unique and we refer to it as the generator.
A.6.2 Given polynomials Ρ'·, j = 1,...,/, any ideal that contains
them all must contain all the polynomials Ρ = Σ#,Λ ^vith arbitrary
polynomial coefficients q·. On the other hand, the set of all theses
sums is clearly an ideal in F[*]. It follows that the ideal generated by
{Pj} is equal to the set of polynomials of the form Ρ = IL^jPj with
polynomial coefficients q·.
The generator G of this ideal divides every one of the P-'s, and,
since G can be expressed as Σ#/Ρ/, every common factor of all the P,'s
divides G. In other words, G = gcd{P{,... ,PZ}, the greatest common
divisor of {Pj}> This implies
Theorem. Given polynomials Ρ,, j = 1,..., /, there exist polynomials
qj such that gcd{P],... ,Pj = 1l4jPj-
In particular:
Corollary. If Ρχ and P2 are relatively prime, there exist polynomials
q]f q2 such that Pxqx +P24i = 1·
A.6.3 Factorization. A polynomial Ρ in F[*] is irreducible or prime
if it has no proper factors, that is, if every factor of Ρ is either scalar
multiple of Ρ or a scalar.
Lemma. //gcd(P,P1) = 1 andP\PxP2, thenP\P2.
PROOF: There exist q,qx such that qP + qxPx = 1. Then
qPP2 + qxPxP2=P2,
the left-hand side is divisible by Ρ and hence so is P2. <
204
Appendix
An immediate extension of the lemma gives that if an irreducible
polynomial Ρ divides a product of polynomials, then it divides (at least) one
of the factors.
Corollary. Assume
η m
(A.6.2) ΓΉ = ΠΦ*
7=1 k=\
with all the factors irreducible and, for simplicity, monic. Then the
sets of factors {Фк} and {РЛ, including repetitions, are identical.
In fact, every P- divides (at least) one of the Φ^/s and is therefore
equal to one of them. Eliminating equal factors on both sides exhausts
both sets of factors.
Theorem (Prime power factorization). Every PeF[i] admits a
factorization Ρ = ΠΦ™;> where each factor Φ is irreducible in ¥[x), and
they are all distinct.
The factorization is unique up to the order in which the factors are
enumerated, and up to multiplication by non-zero scalars.
PROOF: The uniqueness was established in the corollary above. We
prove the existence of the factorization by induction on deg(P). A
linear polynomial Ρ is irreducible and has the trivial factorization: Ρ =
P. Assume that the statement is valid for all к < η and let Ρ G F[i] of
degree n. If Ρ is irreducible, the factorization is again trivial, Ρ = P.
If Pis reducible, Ρ = ΡχΡ2 where both factors have positive degrees kx
and k2 respectively, and kx + k2 = n. Since k- < n, we can apply the
induction hypothesis, write Ρχ = ΓΊΦ· and ^2 = ΠΨ™;> an(l combine
the two factorizations to a factorization for P. м
A.6.4 A polynomial Ρ Ε F[i] splits over F if Ρ is a product of linear
factors in F[jc]. As we observed in A.6.1, (χ — λ) divides Ρ if and only
if Я is a root of P, i.e., if Ρ(λ) = 0. Thus, finding the linear factors of
Ρ is the same as finding the roots of P.
A.6. Polynomials
205
Lemma. //Ψ G F[i] is irreducible, then there is a finite extension of
¥ in which Ψ has a root.
PROOF: Let A be the companion matrix of Ψ. From the discussion in
5.3.3, it follows that Ψ = minPA. Since Ψ is irreducible, the algebra
^(A) = {P(A) : Ρ G ¥[x}} is a field (Proposition 5.3.6). The subfield
{al: a G F} of 3^{Α) is isomorphic to F, so 3^{A) is a finite extension
of F (of degree = deg1!1). Since Ψ(Α) = minPA(A) = 0, Ψ has a root
in &>(A). 4
If Φ G F[jc] does not split over F, then Φ has an irreducible,
nonlinear factor Ψj G ¥[x]. By the lemma, there is a finite extension Kx of
F in which x¥l has a root. If Φ does not split over Kl5 then Φ has an
irreducible, nonlinear factor Ψ2 G Kx[x]. Applying the lemma again,
we find a finite extension K2 of Kx in which Ψ2 has a root. Repeating
this process no more than degΦ times proves the following basic result
from field theory.
Theorem. 1/Фе¥[х], then ¥ has a finite extension К such that Φ
splits in K[x).
Examples:
a. The polynomial x2 + 1 is irreducible in R[x] and splits in C[x].
b. The polynomial χ2 + χ + 1 is irreducible in Z2 [x] and splits in F4 [x].
A.6.5 The fundamental theorem of algebra.
Definition: A field F is algebraically closed if every Ρ e ¥[x] has
roots in F, that is, elements λ G F such that Ρ (λ) = 0.
The fundamental theorem of algebra states that С is algebraically
closed.
Theorem. Given α non-constant polynomial Ρ with complex
coefficients, there exists a complex number λ such that Ρ (λ) = 0.
206
Appendix
A.6.6 It is an immediate consequence of the definitions that F is
algebraically closed if and only if every polynomial in F[jc] splits in F[jc],
or equivalently, if and only if the only irreducible (prime) polynomials
in F[jc] are linear.
Thus, over an algebraically closed field, the prime-power
factorization theorem takes the form:
Theorem. Let F be algebraically closed and let Ρ £F[x)be a
polynomial of degree n. There exist λχ,..., λη £ F (not necessarily distinct)
and αφ§ {the leading coefficient ofP) such that
η
(A.6.3) Ρ(χ) = ά[\(χ-λ]).
ι
A.6.7 Factorization in R[x]. The factorization (A.6.3) applies, of
course, to polynomials with real coefficients, but the roots need not be
real. The basic example is P(x) = x2 + 1 with the roots ±i.
We observe that for polynomials Ρ whose coefficients are all real,
we have Ρ (λ) = Ρ (λ), which means in particular that if Я is a root of
P, then so is A.
A second observation is that
(A.6.4) (χ-λ)(χ-λ)=χ2-2χΚλ + \λ\2
has real coefficients.
Combining these observations with (A.6.3) we obtain that the prime
factors in R[x] are linear polynomials and quadratic polynomials of the
form (A.6.4) where λ £ R.
Theorem. Let Ρ £ R[x] be a polynomial of degree η. Ρ admits a
factorization
(A.6.5) P(z) = aYl(x-Xj)Y\Qj(x),
where a is the leading coefficient, {A ·} is the set of real zeros of Ρ and
<2 are irreducible quadratic polynomials of the form (A.6.4)
corresponding to (pairs of conjugate) nonreal roots of P.
A.6. Polynomials
207
Either product may be empty, in which case it is interpreted as 1.
As mentioned above, the factors appearing in (A.6.5) need not
be distinct—the same factor may be repeated several times. We can
rewrite the product as
(A.6.6) P(z) = a Y\(x - Xj)lJ Π φ (χ),
with A · and Q- now distinct, and the exponents / resp. k- their
multiplicities. The factors {χ — λ·)1} and Q.j(x) appearing in (A.6.6) are
^ j
pairwise relatively prime.
A.6.8 The symmetric functions theorem.
Definition: A polynomial in m variables P{xx,... ,xm) is symmetric
if, for any permutation τ G Sw,
(A.6.7) Ρ(χτ(1)>· ' · >*т(т)) = P(XV ' · · Λ)·
Examples:
b. Gk = ok{xv... ,xm) = £г <...<Ι· χ{ "·χ{ . We assume that к < т in
this case.
The Symmetric Functions Theorem states that every symmetric
polynomial in m variables of degree < к is a polynomial, with rational
coefficients, in {jj,... ,sk}. For our purposes, (see 5.1.3), the following
special case is sufficient.
Proposition. ok(xx,... ,xm) is a polynomial with rational coefficients
in the functions sl(xl,.. .,*m),.. .,sk(xx,... ,xm).
PROOF: Observe that σχ = sx, and s\ = s2 + 2σ2, so σ2 = ^(^ — s2).
In general, we observe that s\—sk — k\ak is a polynomial (with integer
coefficients) in {sp···,^^} and {σ1?... ,σ^}, and the statement
follows by induction. м
208
Appendix
Corollary. Let Ρ be a monic polynomial of degree η in F[jc], and let
{Aj,..., A«} be the roots ofP, repeated according to their multiplicity,
and lying (perhaps) in some finite extension of¥. Then the coefficients
of Ρ are polynomials with rational coefficients in
51(А1,...,Л71),...,5/1_1(А1,...,Л71).
Proof: We have
7=0
and the statement follows from the proposition above. м
*A.6.9 Continuous dependence of the zeros of a polynomial on its
coefficients. Let P(z) = zn + Lq~ * a\^ be a monic polynomial and let
r > LoWj\ = 1 +Σο-1Ι*,·Ι· If Ы > r9 then \z\n > |1о-1я/|, so that
P(z) φ 0. All the zeros of Ρ are located in the disc {z : \z\ < r}.
Denote Ε = {A^} = {z : P(z) = 0}, the set of zeros of P, and, for
η > 0, denote by Εη the "η-neighborhood" of £, that is, the set of
points whose distance from Ε is less than η.
The key remark is that \P(z)\ is bounded away from zero in the
complement of Εη for every η > 0. In other words, given η > 0, there
is a positive ε such that the set {z : \P(z)\ < ε} is contained in Εη.
Proposition. With the preceding notation, given η > 0, there exists
δ > 0 such that ifPx(z) = Lo^/^7 an(^ \a\ ~^j\ < ^, then all the zeros
ofPx are contained in E§.
PROOF: If \P{z)\ > ε on the complement of Εη, take δ small enough
to guarantee that \P{ (z) - P(z) | < § in |z| < 2r. A
Corollary. Let A G JZ(n,C) and η > 0 be given. There exists δ > 0
such that ifAl G <y#(n,C) has all its entries within δ from the
corresponding entries of A, then the spectrum of Αχ is contained in the
η-neighborhood of the spectrum of A.
A.6. Polynomials
209
PROOF: The closeness condition on the entries of Ax to those of A
implies that the coefficients of χΑ are within δ' from those of χΑ, and
δ' can be guaranteed to be arbitrarily small by taking δ small enough.
This page intentionally left blank
Index
Adjoint
of a matrix, 112
of an operator, 63, 112
Algebra, 200
Alternating form, 72
Annihilator, 59
Automorphism, 7
Axiom of choice, 20
Basic partition, 175
Basis, 15
dual, 58
standard, 17
Bilinear
form, 59, 63, 69
map, 69
Canonical
prime-power decomposition,
140
Cauchy-Schwarz, 105
Cayley-Hamilton, 95
Character, 124
Characteristic polynomial
of a matrix, 86
of an operator, 85
Codimension, 18
Cofactor, 80
Complement, 11
Composition , 39
Congruent
matrices, 163
Conjugate
matrices, 43
operators, 40
Conjugation
group actions, 196
in S„, 67
Connecting chain, 171
Contraction, 132
Coset, 10, 192
Cycle, 65, 66
Cyclic
decomposition, 152
system, 94
vector, 94
Decomposition
cyclic, 150, 152
general], 138
prime power, 140
Degrees of freedom, 28
Determinant
of a matrix, 79
of an operator, 76
Diagonal matrix, 44
Diagonal sum, 81, 136
Diagonalizable, 49
Dimension, 17, 18
Direct sum
formal, 10
of subspaces, 11
Eigenspace, 86
generalized, 142
Eigenvalue, 64, 86, 89
Eigenvector, 64, 86, 89
dominant, 168
211
212
Index
Elementary divisors, 155
Equivalence relation, 187
Euclidean space, 103
Factorization
in KM, 206
prime-power, 140, 141, 204
Field, 2
extension, 200
Fixed point, 65
Flag, 92
Fourier expansion, 125
Frobenius, 175
Gauss-Jordan elimination, 26
Gaussian elimination, 24
Group, 1, 189
abelian, 1
cyclic, 192
dual, 125
general linear, 40, 43
Hadamard's inequality, 110
Hamel basis, 20
Hermitian
form, 103, 163
Ideal, 197
Idempotent, 41, 110
Independent
subspaces, 11, 13
vectors, 14
Inertia, law of, 165
Inner product, 103
Irreducible
polynomial, 203
system, 135
Isomorphism, 6
Jordan canonical form, 151, 156
k-form, 69
Kernel, 51
Lagrange, 62
Linear
combination, 9
operator, 35
system, 40, 49
Linear equations
homogeneous, 22
nonhomogeneous, 22
Markov chain, 178
reversible, 179
Matrix
augmented, 25
companion, 96
derogatory, 97
diagonal, 9
Hermitian, 113
integral, 82
nonderogatory, 97
nonnegative, 166, 170
orthogonal, 122
permutation, 44
positive, 167
self-adjoint, 113
skew-symmetric, 9
stochastic, 178
strongly transitive, 175
symmetric, 8
transitive, 171
triangular, 9, 81,92
unimodular, 82
unipotent, 151
unitary , 122
Minimal
system, 100
Minimal polynomial, 96
for(7»,93
Index
213
Minmax principle, 118
Monic polynomial, 201
Multilinear
form, 69
map, 69
Nilpotent, 147
Nilspace, 142
Norm
of an operator, 56
on a vector space, 32
Normal
operator, 119
Nullity, 51
Nullspace, 52
Operator
derogatory, 97
induced on a quotient, 78
linear, 35
nonderogatory, 97
nonnegative definite, 127
nonsingular, 52
normal, 119
orthogonal, 121
positive definite, 127
self-adjoint, 113
singular, 52
unitary, 121
Order
element, 192
group, 189
Orientation, 77
Orthogonal
operator, 121
projection, 108
vectors, 106
Orthogonal equivalence, 122
Orthonormal, 106
Period group, 174
Periodicity, 174
Permutation, 65, 189
Permutation matrix, 44
Perron, 168
Pivot column, 26
Polar decomposition, 128
Polarization, 111
Primary components, 140
Probability vector, 178
Projection, 41
along a subspace, 36
orthogonal, 108
Quadratic form, 162
positive definite, 165
Quotient space, 9
Range, 51
Rank
column, 29
of a matrix, 29
of an operator, 51
row, 25
Reduced-row-echelon form, 26
Reducing subspace, 135
Regular representation, 185
Representation, 181
equivalent, 182
faithful, 181
reducible, 184
regular, 185
unitary, 181
Restriction of an operator, 78
Return times, 171
Ring, 196
Row equivalence, 25
Schur decomposition, 123
Schur's lemma, 100
214
Self-adjoint
algebra, 117
operator, 113
Semisimple
algebra, 144
system, 142
Shift
Jfc-shift, 148
standard, 148
Similar
matrices, 49
operators, 49
Singular value decomposition, 131
Singular values, 130
Solution-set, 8, 23
Span, 9, 14
Spectral mapping theorem, 89
Spectral norm, 162, 167
Spectral Theorems, 114-120
Spectrum, 89, 141
joint, 117
of a matrix, 86
of an operator, 85
Square-free, 143
Steinitz' lemma, 17
Stochastic, 178
Subalgebra, 200
Subgroup
index, 192
normal, 193
Submatrix, 30
principal, 30
Support, 167
Sylvester matrix, 32
Symmetric form, 72
Symmetric group, 65
Tensor product, 12
Trace, 87
Transition matrix, 178
Index
Transposition, 66
Unitary
group, 183
operator, 121
space, 103
Unitary dilation, 132
Unitary equivalence, 122
Vandermonde, 82
Vector space, 4
complex, 4
real, 4
Symbols
с,з
Q,3
м,з
Τ*, 124
Ζ2,3
1Р,Ъ
<Α, 175
Ασ,44
ΛΓ,46
Α*,36
C([0,l]),6
C~([-l,l]),6
CR([0,1]),6
*г>85
C{X) , 6
сотт[Г], 41
CV!38
CW)V) 48
dim У, 18
F",4
F[x],5
ВД, 8
F[xv...,xk},5
GL(Jf) , 183
GL(y), 40
height[v], 147
Я(Ж(У,Ж0,37
^(У,Ж),37
Ш, 170
^(n;F), 5
Л?(п,т;¥), 5
minPT, 96
minPT v, 94
λΤ(β,Ζ), 82
^af({yj}kj=vW),70
Ж^{У®к), 70
^^ут(У®*),72
Ji^alt{r®k)J2
0{n), 122
π^, 108
&>{T),40,9&
ρ (Λ), 29
Sn, 2, 65
spa η [Ε], 9
spa η [7>], 88
II ||ц> > 167
■^V» 78
Гг ,36
^(/ι), 122
JSf(У), 37
215
This page intentionally left blank
Titles in This Series
44 Yitzhak Katznelson and Yonatan R. Katznelson, A (terse)
introduction to linear algebra, 2008
43 Ilka Agricola and Thomas Friedrich, Elementary geometry, 2008
42 С. Е. Silva, Invitation to ergodic theory, 2007
41 Gary L. Mullen and Carl Mummert, Finite fields and applications,
2007
40 Deguang Han, Keri Kornelson, David Larson, and Eric Weber,
Frames for undergraduates, 2007
39 Alex Iosevich, A view from the top: Analysis, combinatorics and number
theory, 2007
38 B. Fristedt, N. Jain, and N. Krylov, Filtering and prediction: A
primer, 2007
37 Svetlana Katok, p-adic analysis compared with real, 2007
36 Mara D. Neusel, Invariant theory, 2007
35 Jorg BewersdorfF, Galois theory for beginners: A historical perspective,
2006
34 Bruce C. Berndt, Number theory in the spirit of Ramanujan, 2006
33 Rekha R. Thomas, Lectures in geometric combinatorics, 2006
32 Sheldon Katz, Enumerative geometry and string theory, 2006
31 John McCleary, A first course in topology: Continuity and dimension,
2006
30 Serge Tabachnikov, Geometry and billiards, 2005
29 Kristopher Tapp, Matrix groups for undergraduates, 2005
28 Emmanuel Lesigne, Heads or tails: An introduction to limit theorems in
probability, 2005
27 Reinhard Illner, C. Sean Bohun, Samantha McCollum, and Thea
van Roode, Mathematical modelling: A case studies approach, 2005
26 Robert Hardt, Editor, Six themes on variation, 2004
25 S. V. Duzhin and B. D. Chebotarevsky, Transformation groups for
beginners, 2004
24 Bruce M. Landman and Aaron Robertson, Ramsey theory on the
integers, 2004
23 S. K. Lando, Lectures on generating functions, 2003
22 Andreas Arvanitoyeorgos, An introduction to Lie groups and the
geometry of homogeneous spaces, 2003
21 W. J. Kaczor and Μ. Τ. Nowak, Problems in mathematical analysis
III: Integration, 2003
20 Klaus Hulek, Elementary algebraic geometry, 2003
19 A. Shen and N. K. Vereshchagin, Computable functions, 2003
18 V. V. Yaschenko, Editor, Cryptography: An introduction, 2002
17 A. Shen and N. K. Vereshchagin, Basic set theory, 2002
TITLES IN THIS SERIES
16 Wolfgang Kuhnel, Differential geometry: curves - surfaces - manifolds,
second edition, 2006
15 Gerd Fischer, Plane algebraic curves, 2001
14 V. A. Vassiliev, Introduction to topology, 2001
13 Frederick J. Almgren, Jr., Plateau's problem: An invitation to varifold
geometry, 2001
12 W. J. Kaczor and Μ. Τ. Nowak, Problems in mathematical analysis
II: Continuity and differentiation, 2001
11 Michael Mesterton-Gibbons, An introduction to game-theoretic
modelling, 2000
®
10 John Oprea, The mathematics of soap films: Explorations with Maple ,
2000
9 David E. Blair, Inversion theory and conformal mapping, 2000
8 Edward B. Burger, Exploring the number jungle: A journey into
diophantine analysis, 2000
7 Judy L. Walker, Codes and curves, 2000
6 Gerald Tenenbaum and Michel Mendes France, The prime numbers
and their distribution, 2000
5 Alexander Mehlmann, The game's afoot! Game theory in myth and
paradox, 2000
4 W. J. Kaczor and Μ. Τ. Nowak, Problems in mathematical analysis
I: Real numbers, sequences and series, 2000
3 Roger Knobel, An introduction to the mathematical theory of waves,
2000
2 Gregory F. Lawler and Lester N. Coyle, Lectures on contemporary
probability, 1999
1 Charles Radin, Miles of tiles, 1999