The Method of Successive Approximations

Suppose for the moment that a solution y ( x ) is known, which reduces to y 0 when x = x 0 ; this solution evidently satisfies the relation

y ( x ) = y 0 + x 0 x f { t , w ( t ) } d t .

This relation is, in reality, an integral equation,1 involving the dependent variable under the integral sign. Let the function y ( x ) be now regarded as unknown; the integral equation may then be solved by a method of successive approximation in the following manner.

Let x lie in the interval [ x 0 , x 0 + h ] and consider the sequence of functions y 1 ( x ) , y 2 ( x ) , , y n ( x ) defined as follows:

y 1 ( x ) = y 0 + x 0 x f { t , y 0 } d t , y 2 ( x ) = y 0 + x 0 x f { t , y 1 ( t ) } d t , y n ( x ) = y 0 + x 0 x f { t , y n 1 ( t ) } d t .

It will now be proved

(a) that, as n increases indefinitely, the sequence of functions y n ( x ) tends to a limit which is a continuous function of x ,

(b) that the limit-function satisfies the differential equation, and

(c) that the solution thus defined assumes the value y 0 when x = x 0 and is the only continuous solution which does so.

In the first place, it will be proved by induction that, when x lies in the interval considered, | y n ( x ) y 0 | b . Suppose then that | y n 1 ( x ) y 0 | b ; it follows that | f { t , y n 1 ( t ) } | M , and consequently

| y n ( x ) y 0 | x 0 x | f { t , y n 1 ( t ) } | d t M ( x x 0 ) M h b .

But evidently | y 1 ( x ) y 0 | b ; it is therefore true that

| y n ( x ) y 0 | b ,

for all values of n . It follows that f { x , y n ( x ) } M when x 0 < x < x 0 + h .

It will now be proved, in a similar way, that

| y n ( x ) y n 1 ( x ) | < M K n 1 n ! ( x x 0 ) n .

For suppose it to be true that, when x 0 x x 0 + h ,

| y n 1 ( x ) y n 2 ( x ) | < M K n 2 ( n 1 ) ! ( x x 0 ) n 1 ,

then

\begin{aligned} |y_n(x) - y_{n-1}(x)| &\leq \int_{x_0}^{x} |f\{t,\, y_{n-1}(t)\} - f\{t,\, y_{n-2}(t)\}|\,dt\\ &< \int_{x_0}^{x} K\,|y_{n-1}(t) - y_{n-2}(t)|\,dt, \end{aligned}

by the Lipschitz condition, so that

\begin{aligned} |y_n(x) - y_{n-1}(x)| &< \frac{MK^{n-1}}{(n-1)!} \int_{x_0}^{x} |t - x_0|^{n-1}\,dt \\ &= \frac{MK^{n-1}}{n!}\,|x - x_0|^n. \end{aligned}

But the inequality is clearly true when n = 1 , it is therefore true for all values of n . In the same way it can be proved to hold when x 0 h x x 0 , it is therefore true for | x x 0 | h .

It follows that the series

y 0 + r = 1 { y r ( x ) y r 1 ( x ) }

is absolutely and uniformly convergent when | x x 0 | h and moreover each term is a continuous function of x . But

y n ( x ) = y 0 + r = 1 n { y r ( x ) y r 1 ( x ) } ;

consequently the limit-function

y ( x ) = lim n y n ( x )

exists and is a continuous function of x in the interval ( x 0 h , x 0 + h ) .2

Now if it is true that

lim n y n ( x ) = y 0 + lim n x 0 x f { t , y n 1 ( t ) } d t = y 0 + x 0 x lim n f { t , y n 1 ( t ) } d t ,

it will follow that y ( x ) is a solution of the integral equation

y ( x ) = y 0 + x 0 x f { t , y ( t ) } d t .

That the inversion of the order of integration and procedure to the limit is legitimate may be proved as follows:

\begin{aligned} \left|\int_{x_0}^{x} [f\{t,\, y(t)\} - f\{t,\, y_{n-1}(t)\}]\,dt\right| &< K \int_{x_0}^{x} |y(t) - y_{n-1}(t)|\,dt \\ &< K\epsilon_n\,|x - x_0| < K\epsilon_n h, \end{aligned}

where ϵ n is independent of x and tends to zero as n tends to infinity.

The function f { t , y ( t ) } is continuous in the interval x 0 h < t < x 0 + h ; consequently

d y ( x ) d x = d d x x 0 x f { t , y ( t ) } d t = f { x , y ( x ) } .

The limit-function y ( x ) therefore satisfies the differential equation; it also reduces to y 0 when x assumes the value x 0 .

It remains to prove that this solution y ( x ) is unique. Suppose Y ( x ) to be a solution distinct from y ( x ) , satisfying the initial condition Y ( x 0 ) = y 0 , and continuous in an interval (x_0,\, x_0 + h') where h' < h and h' is such that the condition

| Y ( x ) y 0 | < b

is satisfied for this interval. Then, since Y ( x ) is a solution of the given equation, it satisfies the integral equation

Y ( x ) = y 0 + x 0 x f { t , Y ( t ) } d t ,

and consequently

Y ( x ) y n ( x ) = x 0 x [ f { t , Y ( t ) } f { t , y n 1 ( t ) } ] d t .

Let n = 1 , then

Y ( x ) y 1 ( x ) = x 0 x [ f { t , Y ( t ) } f { t , y 0 } ] d t ,

and it follows from the Lipschitz condition that

| Y ( x ) y 1 ( x ) | < K b ( x x 0 ) .

Similarly, when n = 2 ,

\begin{aligned} |Y(x) - y_2(x)| &< \int_{x_0}^{x} |f\{t,\, Y(t)\} - f\{t,\, y_1(t)\}|\,dt \\ &< K \int_{x_0}^{x} |Y(t) - y_1(t)|\,dt \\ &< K \int_{x_0}^{x} Kb(t - x_0)\,dt = \tfrac{1}{2} K^2 b (x - x_0)^2, \end{aligned}

and in general

| Y ( x ) y n ( x ) | < K n b ( x x 0 ) n n ! ,

whence

Y ( x ) = lim n y n ( x ) = y ( x )

for all values of x in the interval (x_0,\, x_0 + h'), and therefore the new solution is identical with the old. There is therefore one and only one continuous solution of the differential equation which satisfies the initial conditions.

3.2.1 Observations on the Method of Successive Approximation

The two main assumptions which were made regarding the behaviour of the function f ( x , y ) in the domain D , namely the assumption of continuity and that of the Lipschitz condition are quite independent of one another. The question arises as to the necessity of these assumptions; it is therefore well to look a little more closely into them and to enquire whether or not they may be unduly restrictive.

In the first place, it will be seen that the continuity of f ( x , y ) is not necessary for the existence of a continuous solution; in fact all that the previous investigation demands is that f ( x , y ) be bounded, and that all integrals of the type

x 0 x | f { t , y n ( t ) } | d t

exist. In particular, f ( x , y ) may admit of a limited number of finite discontinuities.3

Thus, for instance, the differential equation

\begin{aligned} \frac{dy}{dx} &= y(1 - 2x) \text{ when } x > 0, \\ &= -y(2x - 1) \text{ when } x < 0 \end{aligned}

admits of a continuous solution satisfying the initial condition y = 1 when x = 1 . This solution is

\begin{aligned} y &= e^{x - x^2} \text{ when } x \geq 0, \\ &= e^{x^2 - x} \text{ when } x \leq 0, \end{aligned}

and the solution is valid for all real values of x , moreover it is unique.

On the other hand, the Lipschitz condition, or a condition of a similar character, must be imposed in order to ensure the uniqueness of the solution. It is not difficult to construct an equation for which the Lipschitz condition is not satisfied, and which admits of more than one continuous solution fulfilling the initial conditions.4

Thus, for instance, in the equation

d y d x = | y | ,

the Lipschitz condition is violated in any region which includes the line y = 0 . The equation admits of two real continuous solutions satisfying the initial conditions x = 0 , y = 0 , viz.

\begin{aligned} (1^\circ)\quad y&=0\\ (2^\circ) \quad y &= \frac{1}{4}x^2\quad\text{when } x \geq 0,\\ &= -\frac{1}{4}x^2\quad \text{when } x \leq 0. \end{aligned}

Another example is given by the equation

d y d x = f ( x , y ) ,

where

\begin{aligned} f(x, y) &= \frac{4x^3 y}{x^4 + y^2} &&\text{ when } x \text{ and } y \text{ are not both zero}, \\ &= 0 &&\text{ when } x = y = 0. \end{aligned}

It is easily proved that f ( x , y ) is a continuous function of x and y . On the other hand

f ( x , Y ) f ( x , y ) = 4 x 3 ( x 4 y Y ) ( x 4 + y 2 ) ( x 4 + Y 2 ) ( Y y ) .

If y = p x 2 , Y = q x 2 ,

| f ( x , Y ) f ( x , y ) | = 4 | 1 p q ( 1 + p 2 ) ( 1 + q 2 ) | | Y y | | x |

and therefore the Lipschitz condition is not satisfied throughout any region containing the origin.

The equation admits of the solution

y = c 2 ( x 4 + c 4 ) ,

c being an arbitrary real constant, and thus there is an infinity of solutions satisfying the initial conditions x = 0 , y = 0 .

The question has been placed on a firm basis by Osgood,5 who proved that, if f ( x , y ) be continuous in the neighbourhood of ( x 0 , y 0 ) , there exists in general a one-fold infinity of solutions satisfying the initial conditions. These solutions lie entirely within the area bounded by two extremal solutions

y = Y 1 ( x ) , y = Y 2 ( x ) .

A necessary and sufficient condition that there be a unique solution is that Y 1 ( x ) and Y 2 ( x ) be identical. This is the case when the Lipschitz condition is satisfied, but it is also true when the Lipschitz condition is replaced by one or other of the less restrictive conditions

\begin{aligned} |f(x, Y) - f(x, y)| &< K_1\,|Y - y|\,\log\frac{1}{|Y - y|},\\ |f(x, Y) - f(x, y)| &< K_2\,|Y - y|\,\log\frac{1}{|Y - y|}\,\log\log\frac{1}{|Y - y|},\\ &\vdots \end{aligned}

in which K 1 , K 2 , are constants.

The constant K which occurs in the Lipschitz condition determines, for any given value of x , the rapidity with which the comparison series

M K n 1 n ! | x x 0 | n

converges, and therefore gives an indication of the utility of the series

y 0 + r = 1 n { y r ( x ) y r 1 ( x ) }

as an approximation to the limit-function y ( x ) . Thus if K were small, y n ( x ) would tend to the limit y ( x ) more rapidly than if K were large. Now in most cases occurring in practice K is the upper bound of

f ( x , y ) y

in the domain D . To make use of this fact, consider the family of curves

f ( x , y ) = C ,

for all values of the constant C . The typical curve of this family is such that it intersects each integral curve in a point at which the gradient of the latter curve is C . For this reason the curves are known as the isoclinal lines.6 Let the isoclinal lines be plotted for a succession of discrete equally-spaced (e.g. integral) values of C , and let a line be drawn parallel to the y -axis. Then the intervals along this line in which the points of intersection with the isoclinal lines are densely packed correspond to large values of K , whereas those intervals in which the intersections are more widely spaced correspond to smaller values of K . This brings out the fact that the regions in which the method of successive approximations may most successfully be applied as a practical method of computation are those in which the isoclinal lines tend to run more or less parallel to the y -axis.7

The method of successive approximations leads to a solution which was shown to converge in the interval | x x 0 | h , where h is the least of a and b / M . But, as was remarked in passing, the assumption originally made that certain conditions are satisfied throughout the region | x x 0 | a , | y y 0 | b was unnecessarily restrictive. If a region | x x 0 | k , | y y 0 | M | x x 0 | can be found such that f ( x , y ) satisfies the necessary conditions in that region, and M is the upper bound of | f ( x , y ) | , then k will certainly not be less, and may quite conceivably be greater, than h . Several writers have succeeded in thus extending the range in which the solution can be proved to converge, but no general method of determining the exact boundaries of the interval of convergence has yet been discovered.

3.2.2 Variation of the Initial Conditions

Let the given initial condition that y = y 0 when x = x 0 be replaced by the new condition y = y 0 + η when x = x 0 , where ( x 0 , y 0 + η ) is a point within the domain D such that | η | δ . Then, in place of the sequence of functions

y 1 ( x ) , y 2 ( x ) , , y n ( x ) ,

as defined in § 3.2, there now arises the sequence

Y 1 ( x ) , Y 2 ( x ) , , Y n ( x ) ,

defined as follows:

Y 1 ( x ) = y 0 + η + x 0 x f { t , y 0 + η } d t , Y 2 ( x ) = y 0 + η + x 0 x f { t , Y 1 ( t ) } d t , Y n ( x ) = y 0 + η + x 0 x f { t , Y n 1 ( t ) } d t .

The existence and uniqueness of the solution

Y ( x ) = lim n Y n ( x )

then follow as before. Now

\begin{aligned} |Y_1(x) - y_1(x)| &\leq \delta + \int_{x_0}^{x} |f\{t,\, y_0 + \eta\} - f\{t,\, y_0\}|\,dt \\ &< \delta + K\delta\,|x - x_0|, \end{aligned}\begin{aligned} |Y_2(x) - y_2(x)| &\leq \delta + \int_{x_0}^{x} |f\{t,\, Y_1(t)\} - f\{t,\, y_1(t)\}|\,dt \\ &< \delta + K\delta\,|x - x_0| + \tfrac{1}{2}K^2\delta\,|x - x_0|^2, \end{aligned}

and, by induction,

\begin{aligned} |Y_n(x) - y_n(x)| &\leq \delta + K\delta\,|x - x_0| + \ldots + \frac{1}{n!}K^n\delta\,|x - x_0|^n \\ &< \delta e^{K|x - x_0|}, \end{aligned}

so that, in the limit,

| Y ( x ) y ( x ) | δ e K | x x 0 | .

Consequently, when | x x 0 | h , the solution is uniformly continuous in the initial value y 0 . To bring out this fact, it may be written in either of the forms

y ( x , y 0 ) and y ( x x 0 , y 0 ) .

Moreover,

| y n ( x , y 0 + η ) y n ( x , y 0 ) η | < 1 + K | x x 0 | + + 1 n ! K n | x x 0 | n ,

and consequently

| y n ( x , y 0 ) y 0 | 1 + K | x x 0 | + + 1 n ! K n | x x 0 | n ,

from which it may be deduced that the series

y ( x , y 0 ) y 0 = 1 + n = 1 { y n ( x , y 0 ) y n 1 ( x , y 0 ) } y 0

is absolutely and uniformly convergent. Therefore y ( x , y 0 ) is uniformly differentiable with respect to y 0 when | x x 0 | h .

A proof proceeding on similar lines to the above shows that if the differential equation involves a parameter λ , that is to say if

d y d x = f ( x , y ; λ ) ,

where f ( x , y ; λ ) is single-valued and continuous and satisfies the Lipschitz condition uniformly in D when Λ 1 λ Λ 2 , then the solution depends continuously upon λ , and in fact is uniformly differentiable with respect to λ when | x x 0 | h .

3.2.3 Singular Points

A singular point may be defined as a point of the ( x , y ) -plane at which one or other of the conditions necessary for the establishment of the existence theorem ceases to hold. In fact if for the initial value-pair ( x 0 , y 0 ) the solution

(a) is discontinuous, (b) is not unique, or (c) does not exist,

then the point ( x 0 , y 0 ) is a singular point of the equation. As illustrations of the diverse ways in which the solutions of an equation may behave at or in the neighbourhood of a singular point, the following examples may be taken.

( 1 ) d y d x = y x .

The conditions requisite for the existence of a unique and continuous solution are fulfilled except in the neighbourhood of x = 0 . The solution corresponding to the initial value-pair ( x 0 , y 0 ) is

y = y 0 x 0 x ,

when x 0 0 . If x 0 = 0 and y 0 0 , the solution reduces to x = 0 .

The only exceptional case is when x 0 = y 0 = 0 ; the only singular point in the finite part of the ( x , y ) -plane is the origin. Now every integral-curve passes through the origin, which is a node of the integral-curves.

( 2 ) d y d x = m y x .

In this case also, the only singular point is the origin. To any other point ( x 0 , y 0 ) corresponds the solution

y = y 0 ( x x 0 ) m .

The family of integral curves corresponding to all possible values of ( x 0 , y 0 ) touch the x -axis at the origin if m > 1 and the y -axis at the origin if $0 < m < 1$. Thus if m > 0 , every integral-curve passes through the origin.

On the other hand, if m < 0 , say m = p , the solution is

y x p = y 0 x 0 p .

The family of integral-curves is asymptotic to the x - and y -axes. The degenerate curve

y x p = 0

passes through the origin, but no other integral-curve does so. The origin is a saddle-point, for in its neighbourhood the integral-curves resemble the contour lines around a mountain pass.

( 3 ) d y d x = x + y x .

The origin is the only singular point; to any other point ( x 0 , y 0 ) corresponds the solution

y = y 0 x 0 x + x log | x x 0 | .

The origin is a node of the integral-curves.

( 4 ) d y d x = x y .

The solution is, in general,

x 2 + y 2 = x 0 2 + y 0 2 .

No real integral-curve, except the degenerate curve x 2 + y 2 = 0 passes through the origin, which is a focal point.

( 5 ) d y d x = x + y x y .

This equation is most effectively dealt with by means of a transformation to polar co-ordinates

x = r cos θ , y = r sin θ .

It then becomes

d r d θ = r ;

the integral-curves are the family of logarithmic spirals

r = c e θ .

One curve of the family goes through each point of the plane except the origin. No integral-curve passes through the origin, which is a focal point of every curve of the family.

It will be noticed that all these examples are particular cases of the general form

d y d x = a x + b y c x + d y ,

which may be integrated by the method of § 2.12. It will be found that, from the point of view of the behaviour of the integral-curves in the neighbourhood of the origin, the equation is of one or other of three main types according as

    I. ( b c ) 2 + 4 a d > 0 ,

    II. ( b c ) 2 + 4 a d < 0 ,

    III. ( b c ) 2 + 4 a d = 0

In Case I the origin is a node if a d b c < 0 , and a saddle-point if a d b c > 0 ; in Case II the origin is a focal point, and in Case III a node.

Footnotes

  1. Bôcher, Introduction to the Theory of Integral Equations; Whittaker and Watson, Modern Analysis, Chap. XI.

  2. Bromwich, Theory of Infinite Series, § 45.

  3. These may be discrete points or lines parallel to the y -axis; any other lines of discontinuity implying a violation of the Lipschitz condition throughout an interval of finite dimensions. Mie, Math. Ann. 43 (1893), p. 553, has shown that solutions exist whenever f ( x , y ) is continuous in x and discontinuous but integrable (in Riemann's sense) with respect to x .

  4. Peano, Math. Ann. 37 (1890), p. 182; Mie, loc. cit., ante; Perron, Math. Ann. 76 (1915), p. 471.

  5. Monatsh. Math. Phys. 9 (1898), p. 331.

  6. The term is due to Chrystal, see Wedderburn, Proc. Roy. Soc. Edin. 24 (1902), p. 400.

  7. Practical methods of approximate computation based upon the method of successive approximations have been devised by Severini, Rend. Ist. Lombard. (2) 31 (1898), pp. 657, 950; Cotton, C. R. Acad. Sc. Paris, 140 (1905), p. 494; 141 (1905), p. 177; 146 (1908), pp. 274, 510; Math. Ann. 81 (1908), p. 107.