The cell below contains all the paper in raw format. This is useful for the AI to have full context of the paper, while you read through it incrementally.
# Programming as Theory Building
Peter Naur
----------
> Peter Naur's classic 1985 essay "Programming as Theory Building" argues that
> a program is not its source code. A program is a shared mental
> construct (he uses the word theory) that lives in the minds of the people who
> work on it. If you lose the people, you lose the program. The code is merely a
> written representation of the program, and it's lossy, so you can't reconstruct
> a program from its code.
### Introduction
The present discussion is a contribution to the understanding of what
programming is. It suggests that programming properly should be regarded as an
activity by which the programmers form or achieve a certain kind of insight, a
theory, of the matters at hand. This suggestion is in contrast to what appears
to be a more common notion, that programming should be regarded as a production
of a program and certain other texts.
Some of the background of the views presented here is to be found in certain
observations of what actually happens to programs and the teams of programmers
dealing with them, particularly in situations arising from unexpected and
perhaps erroneous program executions or reactions, and on the occasion of
modifications of programs. The difficulty of accommodating such observations in
a production view of programming suggests that this view is misleading. The
theory building view is presented as an alternative.
A more general background of the presentation is a conviction that it is
important to have an appropriate understanding of what programming is. If our
understanding is inappropriate we will misunderstand the difficulties that
arise in the activity and our attempts to overcome them will give rise to
conflicts and frustrations.
In the present discussion some of the crucial background experience will first
be outlined. This is followed by an explanation of a theory of what programming
is, denoted the Theory Building View. The subsequent sections enter into some
of the consequences of the Theory Building View.
### Programming and the Programmers’ Knowledge
I shall use the word programming to denote the whole activity of design and
implementation of programmed solutions. What I am concerned with is the
activity of matching some significant part and aspect of an activity in the
real world to the formal symbol manipulation that can be done by a program
running on a computer. With such a notion it follows directly that the
programming activity I am talking about must include the development in time
corresponding to the changes taking place in the real world activity being
matched by the program execution, in other words program modifications.
One way of stating the main point I want to make is that programming in this
sense primarily must be the programmers’ building up knowledge of a certain
kind, knowledge taken to be basically the programmers’ immediate possession,
any documentation being an auxiliary, secondary product.
As a background of the further elaboration of this view given in the following
sections, the remainder of the present section will describe some real
experience of dealing with large programs that has seemed to me more and more
significant as I have pondered over the problems. In either case the experience
is my own or has been communicated to me by persons having first hand contact
with the activity in question.
Case 1 concerns a compiler. It has been developed by a group A for a Language L
and worked very well on computer X. Now another group B has the task to write a
compiler for a language L + M, a modest extension of L, for computer Y. Group B
decides that the compiler for L developed by group A will be a good starting
point for their design, and get a contract with group A that they will get
support in the form of full documentation, including annotated program texts
and much additional written design discussion, and also personal advice. The
arrangement was effective and group B managed to develop the compiler they
wanted. In the present context the significant issue is the importance of the
personal advice from group A in the matters that concerned how to implement the
extensions M to the language. During the design phase group B made suggestions
for the manner in which the extensions should be accommodated and submitted
them to group A for review. In several major cases it turned out that the
solutions suggested by group B were found by group A to make no use of the
facilities that were not only inherent in the structure of the existing
compiler but were discussed at length in its documentation, and to be based
instead on additions to that structure in the form of patches that effectively
destroyed its power and simplicity. The members of group A were able to spot
these cases instantly and could propose simple and effective solutions, framed
entirely within the existing structure. This is an example of how the full
program text and additional documentation is insufficient in conveying to even
the highly motivated group B the deeper insight into the design, that theory
which is immediately present to the members of group A.
In the years following these events the compiler developed by group B was taken
over by other programmers of the same organization, without guidance from group
A. Information obtained by a member of group A about the compiler resulting
from the further modification of it after about 10 years made it clear that at
that later stage the original powerful structure was still visible, but made
entirely ineffective by amorphous additions of many different kinds. Thus,
again, the program text and its documentation has proved insufficient as a
carrier of some of the most important design ideas.
Case 2 concerns the installation and fault diagnosis of a large real–time
system for monitoring industrial production activities. The system is marketed
by its producer, each delivery of the system being adapted individually to its
specific environment of sensors and display devices. The size of the program
delivered in each installation is of the order of 200,000 lines. The relevant
experience from the way this kind of system is handled concerns the role and
manner of work of the group of installation and fault finding programmers. The
facts are, first that these programmers have been closely concerned with the
system as a full time occupation over a period of several years, from the time
the system was under design. Second, when diagnosing a fault these programmers
rely almost exclusively on their ready knowledge of the system and the
annotated program text, and are unable to conceive of any kind of additional
documentation that would be useful to them. Third, other programmers’ groups
who are responsible for the operation of particular installations of the
system, and thus receive documentation of the system and full guidance on its
use from the producer’s staff, regularly encounter difficulties that upon
consultation with the producer’s installation and fault finding programmer
are traced to inadequate understanding of the existing documentation, but which
can be cleared up easily by the installation and fault finding programmers.
The conclusion seems inescapable that at least with certain kinds of large
programs, the continued adaption, modification, and correction of errors in
them, is essentially dependent on a certain kind of knowledge possessed by a
group of programmers who are closely and continuously connected with them.
### Ryle’s Notion of Theory
If it is granted that programming must involve, as the essential part, a
building up of the programmers’ knowledge, the next issue is to characterize
that knowledge more closely. What will be considered here is the suggestion
that the programmers’ knowledge properly should be regarded as a theory, in
the sense of Ryle \[1949\]. Very briefly, a person who has or possesses a
theory in this sense knows how to do certain things and in addition can support
the actual doing with explanations, justifications, and answers to queries,
about the activity of concern. It may be noted that Ryle’s notion of theory
appears as an example of what K. Popper \[Popper, and Eccles, 1977\] calls
unembodied World 3 objects and thus has a defensible philosophical standing. In
the present section we shall describe Ryle’s notion of theory in more detail.
Ryle \[1949\] develops his notion of theory as part of his analysis of the
nature of intellectual activity, particularly the manner in which intellectual
activity differs from, and goes beyond, activity that is merely intelligent. In
intelligent behaviour the person displays, not any particular knowledge of
facts, but the ability to do certain things, such as to make and appreciate
jokes, to talk grammatically, or to fish. More particularly, the intelligent
performance is characterized in part by the person’s doing them well,
according to certain criteria, but further displays the person’s ability to
apply the criteria so as to detect and correct lapses, to learn from the
examples of others, and so forth. It may be noted that this notion of
intelligence does not rely on any notion that the intelligent behaviour depends
on the person’s following or adhering to rules, prescriptions, or methods. On
the contrary, the very act of adhering to rules can be done more or less
intelligently; if the exercise of intelligence depended on following rules
there would have to be rules about how to follow rules, and about how to follow
the rules about following rules, etc. in an infinite regress, which is absurd.
What characterizes intellectual activity, over and beyond activity that is
merely intelligent, is the person’s building and having a theory, where
theory is understood as the knowledge a person must have in order not only to
do certain things intelligently but also to explain them, to answer queries
about them, to argue about them, and so forth. A person who has a theory is
prepared to enter into such activities; while building the theory the person is
trying to get it.
The notion of theory in the sense used here applies not only to the elaborate
constructions of specialized fields of enquiry, but equally to activities that
any person who has received education will participate in on certain occasions.
Even quite unambitious activities of everyday life may give rise to people’s
theorizing, for example in planning how to place furniture or how to get to
some place by means of certain means of transportation.
The notion of theory employed here is explicitly not confined to what may be
called the most general or abstract part of the insight. For example, to have
Newton’s theory of mechanics as understood here it is not enough to
understand the central laws, such as that force equals mass times acceleration.
In addition, as described in more detail by Kuhn \[1970, p. 187ff\], the person
having the theory must have an understanding of the manner in which the central
laws apply to certain aspects of reality, so as to be able to recognize and
apply the theory to other similar aspects. A person having Newton’s theory of
mechanics must thus understand how it applies to the motions of pendulums and
the planets, and must be able to recognize similar phenomena in the world, so
as to be able to employ the mathematically expressed rules of the theory
properly.
The dependence of a theory on a grasp of certain kinds of similarity between
situations and events of the real world gives the reason why the knowledge held
by someone who has the theory could not, in principle, be expressed in terms of
rules. In fact, the similarities in question are not, and cannot be, expressed
in terms of criteria, no more than the similarities of many other kinds of
objects, such as human faces, tunes, or tastes of wine, can be thus expressed.
### The Theory To Be Built by the Programmer
In terms of Ryle’s notion of theory, what has to be built by the programmer
is a theory of how certain affairs of the world will be handled by, or
supported by, a computer program. On the Theory Building View of programming
the theory built by the programmers has primacy over such other products as
program texts, user documentation, and additional documentation such as
specifications.
In arguing for the Theory Building View, the basic issue is to show how the
knowledge possessed by the programmer by virtue of his or her having the theory
necessarily, and in an essential manner, transcends that which is recorded in
the documented products. The answers to this issue is that the programmer’s
knowledge transcends that given in documentation in at least three essential
areas:
1) The programmer having the theory of the program can explain how the solution
relates to the affairs of the world that it helps to handle. Such an
explanation will have to be concerned with the manner in which the affairs of
the world, both in their overall characteristics and their details, are, in
some sense, mapped into the program text and into any additional documentation.
Thus the programmer must be able to explain, for each part of the program text
and for each of its overall structural characteristics, what aspect or activity
of the world is matched by it. Conversely, for any aspect or activity of the
world the programmer is able to state its manner of mapping into the program
text. By far the largest part of the world aspects and activities will of
course lie outside the scope of the program text, being irrelevant in the
context. However, the decision that a part of the world is relevant can only be
made by someone who understands the whole world. This understanding must be
contributed by the programmer.
2) The programmer having the theory of the program can explain why each part of
the program is what it is, in other words is able to support the actual program
text with a justification of some sort. The final basis of the justification is
and must always remain the programmer’s direct, intuitive knowledge or
estimate. This holds even where the justification makes use of reasoning,
perhaps with application of design rules, quantitative estimates, comparisons
with alternatives, and such like, the point being that the choice of the
principles and rules, and the decision that they are relevant to the situation
at hand, again must in the final analysis remain a matter of the programmer’s
direct knowledge.
3) The programmer having the theory of the program is able to respond
constructively to any demand for a modification of the program so as to support
the affairs of the world in a new manner. Designing how a modification is best
incorporated into an established program depends on the perception of the
similarity of the new demand with the operational facilities already built into
the program. The kind of similarity that has to be perceived is one between
aspects of the world. It only makes sense to the agent who has knowledge of the
world, that is to the programmer, and cannot be reduced to any limited set of
criteria or rules, for reasons similar to the ones given above why the
justification of the program cannot be thus reduced.
While the discussion of the present section presents some basic arguments for
adopting the Theory Building View of programming, an assessment of the view
should take into account to what extent it may contribute to a coherent
understanding of programming and its problems. Such matters will be discussed
in the following sections.
### Problems and Costs of Program Modifications
A prominent reason for proposing the Theory Building View of programming is the
desire to establish an insight into programming suitable for supporting a sound
understanding of program modifications. This question will therefore be the
first one to be taken up for analysis.
One thing seems to be agreed by everyone, that software will be modified. It is
invariably the case that a program, once in operation, will be felt to be only
part of the answer to the problems at hand. Also the very use of the program
itself will inspire ideas for further useful services that the program ought to
provide. Hence the need for ways to handle modifications.
The question of program modifications is closely tied to that of programming
costs. In the face of a need for a changed manner of operation of the program,
one hopes to achieve a saving of costs by making modifications of an existing
program text, rather than by writing an entirely new program.
The expectation that program modifications at low cost ought to be possible is
one that calls for closer analysis. First it should be noted that such an
expectation cannot be supported by analogy with modifications of other
complicated man–made constructions. Where modifications are occasionally put
into action, for example in the case of buildings, they are well known to be
expensive and in fact complete demolition of the existing building followed by
new construction is often found to be preferable economically. Second, the
expectation of the possibility of low cost program modifications conceivably
finds support in the fact that a program is a text held in a medium allowing
for easy editing. For this support to be valid it must clearly be assumed that
the dominating cost is one of text manipulation. This would agree with a notion
of programming as text production. On the Theory Building View this whole
argument is false. This view gives no support to an expectation that program
modifications at low cost are generally possible.
A further closely related issue is that of program flexibility. In including
flexibility in a program we build into the program certain operational
facilities that are not immediately demanded, but which are likely to turn out
to be useful. Thus a flexible program is able to handle certain classes of
changes of external circumstances without being modified.
It is often stated that programs should be designed to include a lot of
flexibility, so as to be readily adaptable to changing circumstances. Such
advice may be reasonable as far as flexibility that can be easily achieved is
concerned. However, flexibility can in general only be achieved at a
substantial cost. Each item of it has to be designed, including what
circumstances it has to cover and by what kind of parameters it should be
controlled. Then it has to be implemented, tested, and described. This cost is
incurred in achieving a program feature whose usefulness depends entirely on
future events. It must be obvious that built–in program flexibility is no
answer to the general demand for adapting programs to the changing
circumstances of the world.
In a program modification an existing programmed solution has to be changed so
as to cater for a change in the real world activity it has to match. What is
needed in a modification, first of all, is a confrontation of the existing
solution with the demands called for by the desired modification. In this
confrontation the degree and kind of similarity between the capabilities of the
existing solution and the new demands has to be determined. This need for a
determination of similarity brings out the merit of the Theory Building View.
Indeed, precisely in a determination of similarity the shortcoming of any view
of programming that ignores the central requirement for the direct
participation of persons who possess the appropriate insight becomes evident.
The point is that the kind of similarity that has to be recognized is
accessible to the human beings who possess the theory of the program, although
entirely outside the reach of what can be determined by rules, since even the
criteria on which to judge it cannot be formulated. From the insight into the
similarity between the new requirements and those already satisfied by the
program, the programmer is able to design the change of the program text needed
to implement the modification.
In a certain sense there can be no question of a theory modification, only of a
program modification. Indeed, a person having the theory must already be
prepared to respond to the kinds of questions and demands that may give rise to
program modifications. This observation leads to the important conclusion that
the problems of program modification arise from acting on the assumption that
programming consists of program text production, instead of recognizing
programming as an activity of theory building.
On the basis of the Theory Building View the decay of a program text as a
result of modifications made by programmers without a proper grasp of the
underlying theory becomes understandable. As a matter of fact, if viewed merely
as a change of the program text and of the external behaviour of the execution,
a given desired modification may usually be realized in many different ways,
all correct. At the same time, if viewed in relation to the theory of the
program these ways may look very different, some of them perhaps conforming to
that theory or extending it in a natural way, while others may be wholly
inconsistent with that theory, perhaps having the character of unintegrated
patches on the main part of the program. This difference of character of
various changes is one that can only make sense to the programmer who possesses
the theory of the program. At the same time the character of changes made in a
program text is vital to the longer term viability of the program. For a
program to retain its quality it is mandatory that each modification is firmly
grounded in the theory of it. Indeed, the very notion of qualities such as
simplicity and good structure can only be understood in terms of the theory of
the program, since they characterize the actual program text in relation to
such program texts that might have been written to achieve the same execution
behaviour, but which exist only as possibilities in the programmer’s
understanding.
### Program Life, Death, and Revival
A main claim of the Theory Building View of programming is that an essential
part of any program, the theory of it, is something that could not conceivably
be expressed, but is inextricably bound to human beings. It follows that in
describing the state of the program it is important to indicate the extent to
which programmers having its theory remain in charge of it. As a way in which
to emphasize this circumstance one might extend the notion of program building
by notions of program life, death, and revival. The building of the program is
the same as the building of the theory of it by and in the team of programmers.
During the program life a programmer team possessing its theory remains in
active control of the program, and in particular retains control over all
modifications. The death of a program happens when the programmer team
possessing its theory is dissolved. A dead program may continue to be used for
execution in a computer and to produce useful results. The actual state of
death becomes visible when demands for modifications of the program cannot be
intelligently answered. Revival of a program is the rebuilding of its theory by
a new programmer team.
The extended life of a program according to these notions depends on the taking
over by new generations of programmers of the theory of the program. For a new
programmer to come to possess an existing theory of a program it is
insufficient that he or she has the opportunity to become familiar with the
program text and other documentation. What is required is that the new
programmer has the opportunity to work in close contact with the programmers
who already possess the theory, so as to be able to become familiar with the
place of the program in the wider context of the relevant real world situations
and so as to acquire the knowledge of how the program works and how unusual
program reactions and program modifications are handled within the program
theory. This problem of education of new programmers in an existing theory of a
program is quite similar to that of the educational problem of other activities
where the knowledge of how to do certain things dominates over the knowledge
that certain things are the case, such as writing and playing a music
instrument. The most important educational activity is the student’s doing
the relevant things under suitable supervision and guidance. In the case of
programming the activity should include discussions of the relation between the
program and the relevant aspects and activities of the real world, and of the
limits set on the real world matters dealt with by the program.
A very important consequence of the Theory Building View is that program
revival, that is reestablishing the theory of a program merely from the
documentation, is strictly impossible. Lest this consequence may seem
unreasonable it may be noted that the need for revival of an entirely dead
program probably will rarely arise, since it is hardly conceivable that the
revival would be assigned to new programmers without at least some knowledge of
the theory had by the original team. Even so the Theory Building View suggests
strongly that program revival should only be attempted in exceptional
situations and with full awareness that it is at best costly, and may lead to a
revived theory that differs from the one originally had by the program authors
and so may contain discrepancies with the program text.
In preference to program revival, the Theory Building View suggests, the
existing program text should be discarded and the new–formed programmer team
should be given the opportunity to solve the given problem afresh. Such a
procedure is more likely to produce a viable program than program revival, and
at no higher, and possibly lower, cost. The point is that building a theory to
fit and support an existing program text is a difficult, frustrating, and time
consuming activity. The new programmer is likely to feel torn between loyalty
to the existing program text, with whatever obscurities and weaknesses it may
contain, and the new theory that he or she has to build up, and which, for
better or worse, most likely will differ from the original theory behind the
program text.
Similar problems are likely to arise even when a program is kept continuously
alive by an evolving team of programmers, as a result of the differences of
competence and background experience of the individual programmers,
particularly as the team is being kept operational by inevitable replacements
of the individual members.
### Method and Theory Building
Recent years has seen much interest in programming methods. In the present
section some comments will be made on the relation between the Theory Building
View and the notions behind programming methods.
To begin with, what is a programming method? This is not always made clear,
even by authors who recommend a particular method. Here a programming method
will be taken to be a set of work rules for programmers, telling what kind of
things the programmers should do, in what order, which notations or languages
to use, and what kinds of documents to produce at various stages.
In comparing this notion of method with the Theory Building View of
programming, the most important issue is that of actions or operations and
their ordering. A method implies a claim that program development can and
should proceed as a sequence of actions of certain kinds, each action leading
to a particular kind of documented result. In building the theory there can be
no particular sequence of actions, for the reason that a theory held by a
person has no inherent division into parts and no inherent ordering. Rather,
the person possessing a theory will be able to produce presentations of various
sorts on the basis of it, in response to questions or demands.
As to the use of particular kinds of notation or formalization, again this can
only be a secondary issue since the primary item, the theory, is not, and
cannot be, expressed, and so no question of the form of its expression arises.
It follows that on the Theory Building View, for the primary activity of the
programming there can be no right method.
This conclusion may seem to conflict with established opinion, in several ways,
and might thus be taken to be an argument against the Theory Building View. Two
such apparent contradictions shall be taken up here, the first relating to the
importance of method in the pursuit of science, the second concerning the
success of methods as actually used in software development.
The first argument is that software development should be based on scientific
manners, and so should employ procedures similar to scientific methods. The
flaw of this argument is the assumption that there is such a thing as
scientific method and that it is helpful to scientists. This question has been
the subject of much debate in recent years, and the conclusion of such authors
as Feyerabend \[1978\], taking his illustrations from the history of physics,
and Medawar \[1982\], arguing as a biologist, is that the notion of scientific
method as a set of guidelines for the practising scientist is mistaken.
This conclusion is not contradicted by such work as that of Polya \[1954,
1957\] on problem solving. This work takes its illustrations from the field of
mathematics and leads to insight which is also highly relevant to programming.
However, it cannot be claimed to present a method on which to proceed. Rather,
it is a collection of suggestions aiming at stimulating the mental activity of
the problem solver, by pointing out different modes of work that may be applied
in any sequence.
The second argument that may seem to contradict the dismissal of method of the
Theory Building View is that the use of particular methods has been successful,
according to published reports. To this argument it may be answered that a
methodically satisfactory study of the efficacy of programming methods so far
never seems to have been made. Such a study would have to employ the well
established technique of controlled experiments (cf. \[Brooks, 1980\] or
\[Moher and Schneider, 1982\]). The lack of such studies is explainable partly
by the high cost that would undoubtedly be incurred in such investigations if
the results were to be significant, partly by the problems of establishing in
an operational fashion the concepts underlying what is called methods in the
field of program development. Most published reports on such methods merely
describe and recommend certain techniques and procedures, without establishing
their usefulness or efficacy in any systematic way. An elaborate study of five
different methods by C. Floyd and several co–workers \[Floyd, 1984\]
concludes that the notion of methods as systems of rules that in an arbitrary
context and mechanically will lead to good solutions is an illusion. What
remains is the effect of methods in the education of programmers. This
conclusion is entirely compatible with the Theory Building View of programming.
Indeed, on this view the quality of the theory built by the programmer will
depend to a large extent on the programmer’s familiarity with model solutions
of typical problems, with techniques of description and verification, and with
principles of structuring systems consisting of many parts in complicated
interactions. Thus many of the items of concern of methods are relevant to
theory building. Where the Theory Building View departs from that of the
methodologists is on the question of which techniques to use and in what order.
On the Theory Building View this must remain entirely a matter for the
programmer to decide, taking into account the actual problem to be solved.
### Programmers’ Status and the Theory Building View
The areas where the consequences of the Theory Building View contrast most
strikingly with those of the more prevalent current views are those of the
programmers’ personal contribution to the activity and of the programmers’
proper status.
The contrast between the Theory Building View and the more prevalent view of
the programmers’ personal contribution is apparent in much of the common
discussion of programming. As just one example, consider the study of
modifiability of large software systems by Oskarsson \[1982\]. This study gives
extensive information on a considerable number of modifications in one release
of a large commercial system. The description covers the background, substance,
and implementation, of each modification, with particular attention to the
manner in which the program changes are confined to particular program modules.
However, there is no suggestion whatsoever that the implementation of the
modifications might depend on the background of the 500 programmers employed on
the project, such as the length of time they have been working on it, and there
is no indication of the manner in which the design decisions are distributed
among the 500 programmers. Even so the significance of an underlying theory is
admitted indirectly in statements such as that ‘decisions were implemented in
the wrong block’ and in a reference to ‘a philosophy of AXE’. However, by
the manner in which the study is conducted these admissions can only remain
isolated indications.
More generally, much current discussion of programming seems to assume that
programming is similar to industrial production, the programmer being regarded
as a component of that production, a component that has to be controlled by
rules of procedure and which can be replaced easily. Another related view is
that human beings perform best if they act like machines, by following rules,
with a consequent stress on formal modes of expression, which make it possible
to formulate certain arguments in terms of rules of formal manipulation. Such
views agree well with the notion, seemingly common among persons working with
computers, that the human mind works like a computer. At the level of
industrial management these views support treating programmers as workers of
fairly low responsibility, and only brief education.
On the Theory Building View the primary result of the programming activity is
the theory held by the programmers. Since this theory by its very nature is
part of the mental possession of each programmer, it follows that the notion of
the programmer as an easily replaceable component in the program production
activity has to be abandoned. Instead the programmer must be regarded as a
responsible developer and manager of the activity in which the computer is a
part. In order to fill this position he or she must be given a permanent
position, of a status similar to that of other professionals, such as engineers
and lawyers, whose active contributions as employers of enterprises rest on
their intellectual proficiency.
The raising of the status of programmers suggested by the Theory Building View
will have to be supported by a corresponding reorientation of the programmer
education. While skills such as the mastery of notations, data representations,
and data processes, remain important, the primary emphasis would have to turn
in the direction of furthering the understanding and talent for theory
formation. To what extent this can be taught at all must remain an open
question. The most hopeful approach would be to have the student work on
concrete problems under guidance, in an active and constructive environment.
### Conclusions
Accepting program modifications demanded by changing external circumstances to
be an essential part of programming, it is argued that the primary aim of
programming is to have the programmers build a theory of the way the matters at
hand may be supported by the execution of a program. Such a view leads to a
notion of program life that depends on the continued support of the program by
programmers having its theory. Further, on this view the notion of a
programming method, understood as a set of rules of procedure to be followed by
the programmer, is based on invalid assumptions and so has to be rejected. As
further consequences of the view, programmers have to be accorded the status of
responsible, permanent developers and managers of the activity of which the
computer is a part, and their education has to emphasize the exercise of theory
building, side by side with the acquisition of knowledge of data processing and
notations.
### References
Brooks, R. E. Studying programmer behaviour experimentally. Comm. ACM 23(4):
207–213, 1980.
Feyerabend, P. Against Method. London, Verso Editions, 1978; ISBN:
86091–700–2.
Floyd, C. Eine Untersuchung von Software–Entwicklungs–Methoden. Pp.
248–274 in Programmierumgebungen und Compiler, ed H. Morgenbrod and W.
Sammer, Tagung I/1984 des German Chapter of the ACM, Stuttgart, Teubner Verlag,
1984; ISBN: 3–519–02437–3.
Kuhn, T.S. The Structure of Scientific Revolutions, Second Edition. Chicago,
University of Chicago Press, 1970; ISBN: 0–226–45803–2.
Medawar, P. Pluto’s Republic. Oxford, University Press, 1982: ISBN:
0–19–217726–5.
Moher, T., and Schneider, G. M. Methodology and experimental research in
software engineering, Int. J. Man–Mach. Stud. 16: 65–87, 1. Jan. 1982.
Oskarsson, Ö Mechanisms of modifiability in large software systems Linköping
Studies in Science and Technology, Dissertations, no. 77, Linköping, 1982;
ISBN: 91–7372–527–7.
Polya, G. How To Solve It . New York, Doubleday Anchor Book, 1957.
Polya, G. Mathematics and Plausible Reasoning. New Jersey, Princeton University
Press, 1954.
Popper, K. R., and Eccles, J. C. The Self and Its Brain. London, Routledge and
Kegan Paul, 1977.
Ryle, G. The Concept of Mind. Harmondsworth, England, Penguin, 1963, first
published 1949. Applying "Theory Building"
## Applying “Theory Building”
Viewing programming as theory building helps us understand “metaphor
building” activity in Extreme Programming (XP), and the respective roles of
tacit knowledge and documentation in passing along design knowledge.
### The Metaphor as a Theory
Kent Beck suggested that it is useful to a design team to simplify the general
design of a program to match a single metaphor. Examples might be, “This
program really looks like an assembly line, with things getting added to a
chassis along the line,” or “This program really looks like a restaurant,
with waiters and menus, cooks and cashiers.”
If the metaphor is good, the many associations the designers create around the
metaphor turn out to be appropriate to their programming situation.
That is exactly Naur’s idea of passing along a theory of the design.
If “assembly line” is an appropriate metaphor, then later programmers,
considering what they know about assembly lines, will make guesses about the
structure of the software at hand and find that their guesses are “close.”
That is an extraordinary power for just the two words, “assembly line.”
The value of a good metaphor increases with the number of designers. The closer
each person’s guess is “close” to the other people’s guesses, the
greater the resulting consistency in the final system design.
Imagine 10 programmers working as fast as they can, in parallel, each making
design decisions and adding classes as she goes. Each will necessarily develop
her own theory as she goes. As each adds code, the theory that binds their work
becomes less and less coherent, more and more complicated. Not only maintenance
gets harder, but their own work gets harder. The design easily becomes a
“kludge.” If they have a common theory, on the other hand, they add code in
ways that fit together.
An appropriate, shared metaphor lets a person guess accurately where someone
else on the team just added code, and how to fit her new piece in with it.
### Tacit Knowledge and Documentation
The documentation is almost certainly behind the current state of the program,
but people are good at looking around. What should you put into the
documentation?
That which helps the next programmer build an adequate theory of the program.
This is enormously important. The purpose of the documentation is to jog
memories in the reader, set up relevant pathways of thought about experiences
and metaphors.
This sort of documentation is more stable over the life of the program than
just naming the pieces of the system currently in place.
The designers are allowed to use whatever forms of expression are necessary to
set up those relevant pathways. They can even use multiple metaphors, if they
don’t find one that is adequate for the entire program. They might say that
one section implements a fractal compression algorithm, a second is like an
accounting ledger, the user interface follows the model-observer design
pattern, and so on.
Experienced designers often start their documentation with just
- The metaphors
- Text describing the purpose of each major component
- Drawings of the major interactions between the major components
These three items alone take the next team a long way to constructing a useful
theory of the design.
The source code itself serves to communicate a theory to the next programmer.
Simple, consistent naming conventions help the next person build a coherent
theory. When people talk about “clean code,” a large part of what they are
referring to is how easily the reader can build a coherent theory of the system.
Documentation cannot—and so need not—say everything. Its purpose is to help
the next programmer build an accurate theory about the system.
## HN Discussions
- https://news.ycombinator.com/item?id=23375193
- https://news.ycombinator.com/item?id=20487652
- https://news.ycombinator.com/item?id=10833278
- https://news.ycombinator.com/item?id=7491661Programming as Theory Building
Peter Naur
Peter Naur's classic 1985 essay "Programming as Theory Building" argues that a program is not its source code. A program is a shared mental construct (he uses the word theory) that lives in the minds of the people who work on it. If you lose the people, you lose the program. The code is merely a written representation of the program, and it's lossy, so you can't reconstruct a program from its code.
Introduction
The present discussion is a contribution to the understanding of what programming is. It suggests that programming properly should be regarded as an activity by which the programmers form or achieve a certain kind of insight, a theory, of the matters at hand. This suggestion is in contrast to what appears to be a more common notion, that programming should be regarded as a production of a program and certain other texts.
Some of the background of the views presented here is to be found in certain observations of what actually happens to programs and the teams of programmers dealing with them, particularly in situations arising from unexpected and perhaps erroneous program executions or reactions, and on the occasion of modifications of programs. The difficulty of accommodating such observations in a production view of programming suggests that this view is misleading. The theory building view is presented as an alternative.
A more general background of the presentation is a conviction that it is important to have an appropriate understanding of what programming is. If our understanding is inappropriate we will misunderstand the difficulties that arise in the activity and our attempts to overcome them will give rise to conflicts and frustrations.
In the present discussion some of the crucial background experience will first be outlined. This is followed by an explanation of a theory of what programming is, denoted the Theory Building View. The subsequent sections enter into some of the consequences of the Theory Building View.
Programming and the Programmers’ Knowledge
I shall use the word programming to denote the whole activity of design and implementation of programmed solutions. What I am concerned with is the activity of matching some significant part and aspect of an activity in the real world to the formal symbol manipulation that can be done by a program running on a computer. With such a notion it follows directly that the programming activity I am talking about must include the development in time corresponding to the changes taking place in the real world activity being matched by the program execution, in other words program modifications.
One way of stating the main point I want to make is that programming in this sense primarily must be the programmers’ building up knowledge of a certain kind, knowledge taken to be basically the programmers’ immediate possession, any documentation being an auxiliary, secondary product.
As a background of the further elaboration of this view given in the following sections, the remainder of the present section will describe some real experience of dealing with large programs that has seemed to me more and more significant as I have pondered over the problems. In either case the experience is my own or has been communicated to me by persons having first hand contact with the activity in question.
Case 1 concerns a compiler. It has been developed by a group A for a Language L and worked very well on computer X. Now another group B has the task to write a compiler for a language L + M, a modest extension of L, for computer Y. Group B decides that the compiler for L developed by group A will be a good starting point for their design, and get a contract with group A that they will get support in the form of full documentation, including annotated program texts and much additional written design discussion, and also personal advice. The arrangement was effective and group B managed to develop the compiler they wanted. In the present context the significant issue is the importance of the personal advice from group A in the matters that concerned how to implement the extensions M to the language. During the design phase group B made suggestions for the manner in which the extensions should be accommodated and submitted them to group A for review. In several major cases it turned out that the solutions suggested by group B were found by group A to make no use of the facilities that were not only inherent in the structure of the existing compiler but were discussed at length in its documentation, and to be based instead on additions to that structure in the form of patches that effectively destroyed its power and simplicity. The members of group A were able to spot these cases instantly and could propose simple and effective solutions, framed entirely within the existing structure. This is an example of how the full program text and additional documentation is insufficient in conveying to even the highly motivated group B the deeper insight into the design, that theory which is immediately present to the members of group A.
In the years following these events the compiler developed by group B was taken over by other programmers of the same organization, without guidance from group A. Information obtained by a member of group A about the compiler resulting from the further modification of it after about 10 years made it clear that at that later stage the original powerful structure was still visible, but made entirely ineffective by amorphous additions of many different kinds. Thus, again, the program text and its documentation has proved insufficient as a carrier of some of the most important design ideas.
Case 2 concerns the installation and fault diagnosis of a large real–time system for monitoring industrial production activities. The system is marketed by its producer, each delivery of the system being adapted individually to its specific environment of sensors and display devices. The size of the program delivered in each installation is of the order of 200,000 lines. The relevant experience from the way this kind of system is handled concerns the role and manner of work of the group of installation and fault finding programmers. The facts are, first that these programmers have been closely concerned with the system as a full time occupation over a period of several years, from the time the system was under design. Second, when diagnosing a fault these programmers rely almost exclusively on their ready knowledge of the system and the annotated program text, and are unable to conceive of any kind of additional documentation that would be useful to them. Third, other programmers’ groups who are responsible for the operation of particular installations of the system, and thus receive documentation of the system and full guidance on its use from the producer’s staff, regularly encounter difficulties that upon consultation with the producer’s installation and fault finding programmer are traced to inadequate understanding of the existing documentation, but which can be cleared up easily by the installation and fault finding programmers.
The conclusion seems inescapable that at least with certain kinds of large programs, the continued adaption, modification, and correction of errors in them, is essentially dependent on a certain kind of knowledge possessed by a group of programmers who are closely and continuously connected with them.
Ryle’s Notion of Theory
If it is granted that programming must involve, as the essential part, a building up of the programmers’ knowledge, the next issue is to characterize that knowledge more closely. What will be considered here is the suggestion that the programmers’ knowledge properly should be regarded as a theory, in the sense of Ryle 1949. Very briefly, a person who has or possesses a theory in this sense knows how to do certain things and in addition can support the actual doing with explanations, justifications, and answers to queries, about the activity of concern. It may be noted that Ryle’s notion of theory appears as an example of what K. Popper Popper, and Eccles, 1977 calls unembodied World 3 objects and thus has a defensible philosophical standing. In the present section we shall describe Ryle’s notion of theory in more detail.
what is an unembodied World 3 object?
An unembodied World 3 object is an abstract piece of knowledge or thought-content that is not identical with any physical thing.
Popper divided things roughly like this:
- World 1: physical things — books, computers, paper, brains, source-code files
- World 2: private mental experiences — thoughts, feelings, perceptions
- World 3: objective contents of thought — theories, arguments, mathematical proofs, stories, designs, problems
So a World 3 object might be a theory like:
“Natural selection explains biological evolution.”
That theory can be written in a book, taught in a lecture, or remembered by a person — but the theory itself is not the ink, the sound waves, or the brain cells.
Unembodied means it is not fully captured in a physical form like a document. In Naur’s context, the “theory” of a program may live in programmers’ understanding: their grasp of why the system is shaped the way it is, what changes would fit, what changes would damage it, and how to explain those judgments.
So in plainer words:
An unembodied World 3 object is a structured idea, theory, or understanding that exists as intellectual content, even though it is not completely written down or physically represented.
Ryle 1949 develops his notion of theory as part of his analysis of the nature of intellectual activity, particularly the manner in which intellectual activity differs from, and goes beyond, activity that is merely intelligent. In intelligent behaviour the person displays, not any particular knowledge of facts, but the ability to do certain things, such as to make and appreciate jokes, to talk grammatically, or to fish. More particularly, the intelligent performance is characterized in part by the person’s doing them well, according to certain criteria, but further displays the person’s ability to apply the criteria so as to detect and correct lapses, to learn from the examples of others, and so forth. It may be noted that this notion of intelligence does not rely on any notion that the intelligent behaviour depends on the person’s following or adhering to rules, prescriptions, or methods. On the contrary, the very act of adhering to rules can be done more or less intelligently; if the exercise of intelligence depended on following rules there would have to be rules about how to follow rules, and about how to follow the rules about following rules, etc. in an infinite regress, which is absurd.
What characterizes intellectual activity, over and beyond activity that is merely intelligent, is the person’s building and having a theory, where theory is understood as the knowledge a person must have in order not only to do certain things intelligently but also to explain them, to answer queries about them, to argue about them, and so forth. A person who has a theory is prepared to enter into such activities; while building the theory the person is trying to get it.
I find the sentence in bold very interesting. It is basically saying that intellectual activity requires an internal model (which here they call an unembodied world 3 object, aka as theory), and that explanations, queries, etc... derive from inspecting this model rather than the thing itself.
Yes — that’s very close, with one useful nuance.
The bold sentence is saying that intellectual competence is not just being able to perform well, but being able to account for the performance:
- explain why this action is appropriate
- answer “what if?” questions
- justify choices
- criticize bad alternatives
- adapt to new cases
- argue about edge cases
So in Naur’s programming context, a programmer who “has the theory” of a program doesn’t merely know the source code. They have a working understanding of the system’s concepts, constraints, purposes, invariants, and design pressures.
Your phrase “internal model” is a good modern gloss, but Ryle/Naur would probably resist making it sound like a literal hidden diagram in the head. It’s less like a private object you stare at, and more like a structured practical understanding that shows itself in what you can do and say.
A nice distinction is:
- merely intelligent activity: you can do the thing well
- intellectual activity: you can do the thing well and explain, defend, adapt, and reason about it
For programming, that means the theory lets you answer questions like:
“If we add this feature here, will it fit the design or damage it?”
That answer does not come simply from inspecting the code as text. It comes from understanding the system’s deeper organization — the “theory” of the program. The code is evidence and expression of the theory, but not the theory itself.
internal models are nothing more than chunks of neurons firing together so clearly it isn't a nicely delimited hidden diagram, but rather a combination of sensory experiences, reflections and memories.
Yes — exactly. That’s why “internal model” is useful but slightly dangerous if taken too literally.
At the physical level, the programmer’s theory is realized in neurons, habits of attention, remembered examples, bodily familiarity with tools, past debugging episodes, language, diagrams they’ve seen, etc. It is not a clean little object stored somewhere.
But at the explanatory level, it can still make sense to call it a theory, because it behaves like a coherent capability:
- it lets you predict what will happen
- it lets you explain why something is wrong
- it lets you judge whether a proposed change “fits”
- it lets you answer questions not already written in the code or docs
So there are two levels:
- implementation level: distributed neural patterns, memories, experiences, associations
- conceptual/practical level: an organized understanding of the system
This also softens the “inspecting the model” phrasing. It is probably not that someone looks inward at a hidden diagram. Rather, when questioned, they can reconstruct, activate, and articulate relevant parts of their understanding.
For Naur, that is the key thing: the “program” is not exhausted by the text. The code is a public artifact, but the living understanding that makes the code intelligible is distributed through the programmers’ trained capacities.
this kindof of flies in the face of all modern SWE. Most teams are organized in such a way that everyone can be replaced, similarly to "The Wall" concept by Pink Floyd. It makes sense to mitigate the disaster of an employee living by making the public artifact as explicit as possible, make a lot of tests, CI/CD etc...
Yes — and I think Naur’s point is especially sharp because modern SWE tries so hard to make people replaceable.
Modern SWE practices are partly an attempt to externalize fragments of the theory:
- tests encode expected behavior
- types/interfaces encode constraints
- CI/CD encodes operational discipline
- docs/RFCs/ADRs encode rationale
- code review spreads judgment across people
- pairing/onboarding transfers tacit understanding
So they are not useless; they are ways of reducing the damage when someone leaves.
But Naur would probably say: those artifacts still do not contain the whole theory. They preserve signs of understanding, not understanding itself.
So modern SWE has two possible interpretations:
- Wise version: “People leave, so let’s externalize and share as much theory as possible.”
- Naive version: “If the repo has enough tests/docs/process, people are interchangeable.”
Naur is attacking the naive version.
The best teams probably do not eliminate theory-building; they distribute it. They try to avoid one person being the sole carrier of the theory, but they still depend on humans having a living understanding of the system. Tests and CI protect against some failures, but they do not replace design judgment.
because of AI is designed, I get the feeling they will be very good at preserving the signs of understanding but they can fail horribly because:
- We already said the theory can not ever be completed expressed in the docs
- The LLMs are trained on the whole internet, so they probably come with a "hidden" average theory of all programs in the internet.
We could try to mitigate 2. by using forcing the AI to "ground" the responses on the codebase (or docs) but because of 1. we can't never (at least with LLMs it seems) get a proper Theory inside the LLM.
Yes — that feels like a very Naur-ish critique of AI coding.
I’d phrase it like this:
LLMs are excellent at manipulating the public traces of theory, but they do not automatically possess the local theory that made those traces meaningful.
And your point 2 is important. An LLM comes with something like an average internet-software prior:
- “projects usually structure auth like this”
- “React apps usually do state like this”
- “Python APIs usually mean this”
- “this error usually implies that fix”
That prior is often useful! But it can be dangerous because a real codebase may have very specific reasons for not doing the average thing.
Grounding the model in the repo helps, but only partly. It can force the model to attend to:
- existing code
- tests
- docs
- names
- comments
- architecture notes
- commit messages, if available
But Naur’s point bites here: those are still artifacts of theory, not the whole theory. They may not include:
- why a weird design was chosen
- which alternative failed three years ago
- which invariant everyone knows but nobody wrote down
- which customer constraint shaped the module
- which “obvious cleanup” would destroy an important property
So I mostly agree with your conclusion, with one nuance.
I wouldn’t say an LLM can never get a proper theory in any sense. It can build a temporary working approximation of a theory during a session, especially if it can inspect code, run tests, ask questions, and receive corrections.
But it is usually not the same as Naur’s programmer-theory, because Naur’s theory is not just stored information. It is a lived competence built through participation, debugging, modification, argument, and responsibility for consequences.
So the practical distinction might be:
- LLM as code-text manipulator: dangerous if treated as interchangeable senior engineer
- LLM as theory-building assistant: useful if it helps humans expose, test, document, and transfer their understanding
In that sense, AI may be best not as a replacement for theory-building, but as a tool for making theory-building more explicit. It can ask “why is this here?”, propose interpretations, generate tests that reveal assumptions, summarize design rationale, and notice inconsistencies. But the local theory still has to be validated by people who actually understand the system.
The notion of theory in the sense used here applies not only to the elaborate constructions of specialized fields of enquiry, but equally to activities that any person who has received education will participate in on certain occasions. Even quite unambitious activities of everyday life may give rise to people’s theorizing, for example in planning how to place furniture or how to get to some place by means of certain means of transportation.
The notion of theory employed here is explicitly not confined to what may be called the most general or abstract part of the insight. For example, to have Newton’s theory of mechanics as understood here it is not enough to understand the central laws, such as that force equals mass times acceleration. In addition, as described in more detail by Kuhn 1970, p. 187ff, the person having the theory must have an understanding of the manner in which the central laws apply to certain aspects of reality, so as to be able to recognize and apply the theory to other similar aspects. A person having Newton’s theory of mechanics must thus understand how it applies to the motions of pendulums and the planets, and must be able to recognize similar phenomena in the world, so as to be able to employ the mathematically expressed rules of the theory properly.
what is this reference about "Kuhn [1970, p. 187ff]" ?
web_search(type="search", queries="['Naur Programming as Theory Building Ku…", query="Naur Programming as Theory Building Kuhn…")→"Server tool call executed."
{
"id": "ws_0038f39e780c6618016a2015d288208191bd188d6f2b198535",
"name": "web_search",
"args": {
"type": "search",
"queries": "['Naur Programming as Theory Building Kuhn 1970 p. 187ff reference', \"Kuhn 1970 p. 187ff exemplars Newton's theory pendulum planets\"]",
"query": "Naur Programming as Theory Building Kuhn 1970 p. 187ff reference"
},
"result": "Server tool call executed.",
"server": true
}
web_search(type="search", queries="['\"The Structure of Scientific Revolutio…", query="\"The Structure of Scientific Revolutions…")→"Server tool call executed."
{
"id": "ws_0038f39e780c6618016a2015e337448191859df342f96fb0ca",
"name": "web_search",
"args": {
"type": "search",
"queries": "['\"The Structure of Scientific Revolutions\" \"p. 187\" \"exemplars\"', '\"The Structure of Scientific Revolutions\" \"187\" \"exemplar\" \"pendulum\"', '\"Second Thoughts on Paradigms\" Kuhn 1970 p. 187']",
"query": "\"The Structure of Scientific Revolutions\" \"p. 187\" \"exemplars\""
},
"result": "Server tool call executed.",
"server": true
}
web_search(type="open_page", url="https://gwern.net/doc/cs/algorithm/1985-…")→"Server tool call executed."
{
"id": "ws_0038f39e780c6618016a20160b65cc81918da5858d4f14d7d4",
"name": "web_search",
"args": {
"type": "open_page",
"url": "https://gwern.net/doc/cs/algorithm/1985-naur.pdf"
},
"result": "Server tool call executed.",
"server": true
}
web_search(type="find_in_page", pattern="Kuhn", url="https://gwern.net/doc/cs/algorithm/1985-…")→"Server tool call executed."
{
"id": "ws_0038f39e780c6618016a201660e610819186dbcca741acfe98",
"name": "web_search",
"args": {
"type": "find_in_page",
"pattern": "Kuhn",
"url": "https://gwern.net/doc/cs/algorithm/1985-naur.pdf"
},
"result": "Server tool call executed.",
"server": true
}
web_search(type="search", queries="['Stanford Encyclopedia Philosophy Thoma…", query="Stanford Encyclopedia Philosophy Thomas …")→"Server tool call executed."
{
"id": "ws_0038f39e780c6618016a2016b121788191b4874189023f0a81",
"name": "web_search",
"args": {
"type": "search",
"queries": "['Stanford Encyclopedia Philosophy Thomas Kuhn exemplars disciplinary matrix symbolic generalizations values models']",
"query": "Stanford Encyclopedia Philosophy Thomas Kuhn exemplars disciplinary matrix symbolic generalizations values models"
},
"result": "Server tool call executed.",
"server": true
}
web_search(type="search", queries="['University of Chicago Press The Struct…", query="University of Chicago Press The Structur…")→"Server tool call executed."
{
"id": "ws_0038f39e780c6618016a2016be498481919e18b74cf54558bf",
"name": "web_search",
"args": {
"type": "search",
"queries": "['University of Chicago Press The Structure of Scientific Revolutions 2nd edition 1970 Kuhn']",
"query": "University of Chicago Press The Structure of Scientific Revolutions 2nd edition 1970 Kuhn"
},
"result": "Server tool call executed.",
"server": true
}
web_search(type="open_page", url="https://plato.stanford.edu/entries/thoma…")→"Server tool call executed."
{
"id": "ws_0038f39e780c6618016a2016c9fbb881919093179159f214fe",
"name": "web_search",
"args": {
"type": "open_page",
"url": "https://plato.stanford.edu/entries/thomas-kuhn/"
},
"result": "Server tool call executed.",
"server": true
}
Kuhn [1970, p. 187ff] refers to Thomas S. Kuhn’s The Structure of Scientific Revolutions, 2nd edition, 1970. Naur’s own reference list gives that exact citation: Kuhn, University of Chicago Press, 1970.
The ff means “and following pages”. So: page 187 and the next few pages.
The relevant bit is Kuhn’s discussion of paradigms as exemplars. In the 1970 postscript, Kuhn clarifies that scientific understanding is not just knowing abstract laws, but also knowing canonical solved examples and being able to recognize similar cases.
So for Newtonian mechanics:
- abstract rule:
F = ma - exemplar cases: pendulums, inclined planes, planetary motion, falling bodies
- real competence: seeing a new physical situation as “like” one of those cases, then applying the theory appropriately
That’s why Naur cites Kuhn here. Naur wants to say: having a “theory” is not just possessing explicit propositions or rules. It includes trained judgment about which similarities matter.
In programming terms, this maps beautifully:
- abstract principle: “separate concerns”
- exemplar: “this kind of domain object belongs behind this interface”
- real competence: recognizing that a new feature is similar to a previous design situation, and fitting it into the system without damaging the structure
So Kuhn is helping Naur support the claim that theory involves pattern-recognition, example-based judgment, and application, not merely written rules.
Sources: Naur’s essay gives the citation and surrounding passage in the PDF, and Stanford Encyclopedia explains Kuhn’s 1970 clarification of paradigms as exemplars/puzzle-solutions.
- Naur PDF: https://gwern.net/doc/cs/algorithm/1985-naur.pdf
- Stanford Encyclopedia of Philosophy on Kuhn: https://plato.stanford.edu/entries/thomas-kuhn/
The dependence of a theory on a grasp of certain kinds of similarity between situations and events of the real world gives the reason why the knowledge held by someone who has the theory could not, in principle, be expressed in terms of rules. In fact, the similarities in question are not, and cannot be, expressed in terms of criteria, no more than the similarities of many other kinds of objects, such as human faces, tunes, or tastes of wine, can be thus expressed.
In the previous short paragraph, the first sentence reminds me a lot of Yann LeCunn insisting that AI needs a world model. If this theory requires a grasp on similar objects of the real world, understanding the real world then is necessary for a theory of the kind.
The second sentence instead, reminds me so much of "Surfaces and Essences" book, where they show that each word can be totally different for each person depending on which associations the word has. For example, studio might evoke completely different things for someone reminded of his grandpa's attic artist studio) vs. someone reminded of his stressing hard-working period of time in a super small studio in a big city.
Yes — those are two very strong connections.
On the LeCun point: Naur’s “theory” seems to require something like a world model, because the programmer must recognize which situations are relevantly similar. That cannot be done from formal rules alone.
For example, in programming:
“This new feature is like that old payment-flow exception, not like the normal checkout path.”
That judgment depends on understanding the world the program participates in: users, business constraints, failure modes, timing, weird edge cases, organizational habits, etc. The source code alone does not tell you which similarity matters.
So if LeCun says AI needs world models, Naur gives a very software-specific version of why: without a model of the relevant world, the AI may match surface patterns rather than grasping the situation.
Your Surfaces and Essences connection is also excellent. Hofstadter and Sander argue that thinking is deeply analogical: concepts are not crisp definitions, but networks of remembered examples and associations. So “studio” is not just a dictionary entry; it activates a personal cloud of experiences, moods, images, purposes, and analogies.
That maps almost directly onto Naur’s line about faces, tunes, and wine. We recognize similarity without being able to state exact criteria.
And this has a big consequence for programming communication:
- a doc says “cache”
- one programmer hears “performance optimization”
- another hears “consistency hazard”
- another hears “distributed invalidation nightmare”
- another hears “temporary memoization inside this request”
Same word, different experiential category.
So documentation does not simply “transfer meaning.” It gives prompts that activate the reader’s existing concepts. If the reader lacks the right examples, history, and analogies, the words may produce the wrong theory.
That’s why examples, counterexamples, pair programming, debugging stories, design reviews, and “why not this other approach?” explanations matter so much. They don’t just state rules; they help shape the reader’s similarity-judgment.
The Theory To Be Built by the Programmer
In terms of Ryle’s notion of theory, what has to be built by the programmer is a theory of how certain affairs of the world will be handled by, or supported by, a computer program. On the Theory Building View of programming the theory built by the programmers has primacy over such other products as program texts, user documentation, and additional documentation such as specifications.
In arguing for the Theory Building View, the basic issue is to show how the knowledge possessed by the programmer by virtue of his or her having the theory necessarily, and in an essential manner, transcends that which is recorded in the documented products. The answers to this issue is that the programmer’s knowledge transcends that given in documentation in at least three essential areas:
- The programmer having the theory of the program can explain how the solution
relates to the affairs of the world that it helps to handle. Such an explanation will have to be concerned with the manner in which the affairs of the world, both in their overall characteristics and their details, are, in some sense, mapped into the program text and into any additional documentation. Thus the programmer must be able to explain, for each part of the program text and for each of its overall structural characteristics, what aspect or activity of the world is matched by it. Conversely, for any aspect or activity of the world the programmer is able to state its manner of mapping into the program text. By far the largest part of the world aspects and activities will of course lie outside the scope of the program text, being irrelevant in the context. However, the decision that a part of the world is relevant can only be made by someone who understands the whole world. This understanding must be contributed by the programmer.
- The programmer having the theory of the program can explain why each part of
the program is what it is, in other words is able to support the actual program text with a justification of some sort. The final basis of the justification is and must always remain the programmer’s direct, intuitive knowledge or estimate. This holds even where the justification makes use of reasoning, perhaps with application of design rules, quantitative estimates, comparisons with alternatives, and such like, the point being that the choice of the principles and rules, and the decision that they are relevant to the situation at hand, again must in the final analysis remain a matter of the programmer’s direct knowledge.
This second point emphasizes the intuitive knowledge of the program over the program itself. I find this super important because as the codebase grows, code might be introduced that works against the Theory principles. For example, the group B extending writing the extensions M for compiler L without leveraging all the goodness introduced by original group A. This rule again is against most common sense. Code is not the source of truth but rather the Theory built by humans. As such, corrections to the codebase can be jutstified on the basis of this direct knowledge. I think this happens often with Jeremy, because he has this clear Theory of answer.ai libraries but of course he can't keep track of all the code in them. So sometimes his PR reviews or comments are done according to the Theory rather than the codebase and I think this is why he can be so effective in a sense. Because he has a sprawling direct and quick knowledge about the Theory and he doesn't need to fully know all the codebase to express it.
Yes — I think you’ve put your finger on exactly what point 2 is doing, and the group A/B case is the perfect illustration.
The key inversion is: the program text is downstream of the theory, not the other way around. So when code violates the theory, the theory is what tells you the code is wrong — not the other way around. Group B’s patches worked in the sense that the code ran, but group A could see instantly that they were “wrong” because they damaged the structure the theory protects.
Your Jeremy example is a really nice real-world case of point 2, and it highlights something subtle: having the theory lets you make justified judgments without re-deriving them from the full code each time. The justification bottoms out in direct, intuitive knowledge — “this doesn’t fit” — and only afterward gets dressed up in reasons (design rules, comparisons, alternatives).
There's a nuance worth pulling out though. Naur isn't saying the theory-holder is infallible about the code. They might misremember a detail, or a specific function might genuinely have changed. What the theory gives them is authority over what the code should be, not perfect recall of what it currently is. Those are different kinds of knowledge — and interestingly, the theory is the more durable one.
That maps onto your group B story over 10 years: the code drifted, but someone with the original theory could still look at the mess and recognize the "powerful structure still visible" underneath. The theory acts as a reference frame against which code drift becomes visible as drift.
what about these 3 layers of debt introduced recently to talk about AI code? https://martinfowler.com/fragments/2026-04-02.html
Let me look that up, since it's a recent article past my knowledge cutoff.
web_search(query="martinfowler.com three layers of debt AI…")→"Server tool call executed."
{
"id": "srvtoolu_01KRuWCoqgsKcHgN9AZb48eE",
"name": "web_search",
"args": {
"query": "martinfowler.com three layers of debt AI code 2026"
},
"result": "Server tool call executed.",
"server": true
}
This maps beautifully onto everything we've been discussing. The framework comes from Margaret-Anne Storey, and Martin Fowler summarizes it as three layers of system health:
*Technical debt lives in code. It accumulates when implementation decisions compromise future changeability. It limits how systems can change. *Cognitive debt lives in people. It accumulates when shared understanding of the system erodes faster than it is replenished. It limits how teams can reason about change. And *intent debt lives in artifacts. It accumulates when the goals and constraints that should guide the system are poorly captured or maintained.
Here's why this should excite you given where we just landed:
Cognitive debt is almost exactly Naur's "loss of theory." It's the erosion of the living understanding in people's heads. And notice the framework insists it lives in people, not in code — precisely Naur's point that the theory has primacy and cannot be fully recorded in the artifacts.
Intent debt is the formalization of point 2 (the "why each part is what it is"). It's the missing justification — the rationale that, as Naur said, ultimately bottoms out in the programmer's direct knowledge but which teams try to externalize into ADRs, specs, etc.
And the connection to your AI worry from earlier is direct: *AI may reduce technical debt while accelerating cognitive and intent debt. When developers write code from scratch, the friction of implementation builds understanding along the way. When AI generates that code, the developer may accept it without building the same mental model. They even have a name for your "preserving the signs without the theory" idea — *the paper calls this "cognitive surrender" — adopting AI outputs with minimal scrutiny.
- The programmer having the theory of the program is able to respond
constructively to any demand for a modification of the program so as to support the affairs of the world in a new manner. Designing how a modification is best incorporated into an established program depends on the perception of the similarity of the new demand with the operational facilities already built into the program. The kind of similarity that has to be perceived is one between aspects of the world. It only makes sense to the agent who has knowledge of the world, that is to the programmer, and cannot be reduced to any limited set of criteria or rules, for reasons similar to the ones given above why the justification of the program cannot be thus reduced.
While the discussion of the present section presents some basic arguments for adopting the Theory Building View of programming, an assessment of the view should take into account to what extent it may contribute to a coherent understanding of programming and its problems. Such matters will be discussed in the following sections.
Problems and Costs of Program Modifications
A prominent reason for proposing the Theory Building View of programming is the desire to establish an insight into programming suitable for supporting a sound understanding of program modifications. This question will therefore be the first one to be taken up for analysis.
One thing seems to be agreed by everyone, that software will be modified. It is invariably the case that a program, once in operation, will be felt to be only part of the answer to the problems at hand. Also the very use of the program itself will inspire ideas for further useful services that the program ought to provide. Hence the need for ways to handle modifications.
The question of program modifications is closely tied to that of programming costs. In the face of a need for a changed manner of operation of the program, one hopes to achieve a saving of costs by making modifications of an existing program text, rather than by writing an entirely new program.
The expectation that program modifications at low cost ought to be possible is one that calls for closer analysis. First it should be noted that such an expectation cannot be supported by analogy with modifications of other complicated man–made constructions. Where modifications are occasionally put into action, for example in the case of buildings, they are well known to be expensive and in fact complete demolition of the existing building followed by new construction is often found to be preferable economically. Second, the expectation of the possibility of low cost program modifications conceivably finds support in the fact that a program is a text held in a medium allowing for easy editing. For this support to be valid it must clearly be assumed that the dominating cost is one of text manipulation. This would agree with a notion of programming as text production. On the Theory Building View this whole argument is false. This view gives no support to an expectation that program modifications at low cost are generally possible.
once again, the idea that low cost modifications are possible is based on programming as text production, which in turn is what LLMs are good at, what they are marketed as, and where the whole industry seems to go (layoffs because AI can't code much faster)
Yes — and notice how sharp Naur's wording is: he says on the Theory Building View "this whole argument is false." Not incomplete, not risky — false. The cheap-modification expectation rests entirely on the premise that the dominating cost is text manipulation. Kill that premise and the whole economic argument collapses.
Which is exactly the premise the current wave is built on. If programming is text production, then a machine that produces text fast and cheap should produce software fast and cheap. The layoffs-because-AI logic only follows if you've already accepted the production view.
But there's a nice irony here that I think strengthens your point: the faster and cheaper the text production becomes, the more the bottleneck shifts to the part Naur said was never about text in the first place — the theory. If anything, cheap text generation makes theory the scarce resource, not the abundant one.
So one way to frame it: AI doesn't refute Naur, it runs his experiment at scale. We're about to find out empirically whether the dominating cost was text or theory.
imagine next billion dollars companies (or any proxy for success) are fully built on code generated by AI, what would that say about Naur's theory
That would be a genuinely interesting challenge to Naur — but I'd want to be careful about which claim it refutes, because Naur is making a narrower point than "AI can't build successful software."
Let me offer a few possible interpretations, and I'm curious which one you find most likely:
1. It would refute Naur cleanly. If billion-dollar systems are built and successfully modified over years with no human holding the theory, then theory wasn't essential after all — the artifacts were enough. Naur was simply wrong about what programs fundamentally are.
2. It wouldn't refute Naur — the theory just moved. Maybe the AI (or the human + AI loop) is now the theory-holder. Naur never said the theory had to live in a human; he said it couldn't live in the text. If something can answer "will this change fit or damage the design?", it has the theory, regardless of substrate.
3. It refutes Naur's scope, not his core. Maybe theory matters intensely for some programs (long-lived, evolving, deeply embedded in a messy world) but barely for others (CRUD apps, glue code, throwaway tools). Billion-dollar success might just prove most software lives in the "theory-cheap" zone.
A further closely related issue is that of program flexibility. In including flexibility in a program we build into the program certain operational facilities that are not immediately demanded, but which are likely to turn out to be useful. Thus a flexible program is able to handle certain classes of changes of external circumstances without being modified.
It is often stated that programs should be designed to include a lot of flexibility, so as to be readily adaptable to changing circumstances. Such advice may be reasonable as far as flexibility that can be easily achieved is concerned. However, flexibility can in general only be achieved at a substantial cost. Each item of it has to be designed, including what circumstances it has to cover and by what kind of parameters it should be controlled. Then it has to be implemented, tested, and described. This cost is incurred in achieving a program feature whose usefulness depends entirely on future events. It must be obvious that built–in program flexibility is no answer to the general demand for adapting programs to the changing circumstances of the world.
In a program modification an existing programmed solution has to be changed so as to cater for a change in the real world activity it has to match. What is needed in a modification, first of all, is a confrontation of the existing solution with the demands called for by the desired modification. In this confrontation the degree and kind of similarity between the capabilities of the existing solution and the new demands has to be determined. This need for a determination of similarity brings out the merit of the Theory Building View. Indeed, precisely in a determination of similarity the shortcoming of any view of programming that ignores the central requirement for the direct participation of persons who possess the appropriate insight becomes evident. The point is that the kind of similarity that has to be recognized is accessible to the human beings who possess the theory of the program, although entirely outside the reach of what can be determined by rules, since even the criteria on which to judge it cannot be formulated. From the insight into the similarity between the new requirements and those already satisfied by the program, the programmer is able to design the change of the program text needed to implement the modification.
In a certain sense there can be no question of a theory modification, only of a program modification. Indeed, a person having the theory must already be prepared to respond to the kinds of questions and demands that may give rise to program modifications. This observation leads to the important conclusion that the problems of program modification arise from acting on the assumption that programming consists of program text production, instead of recognizing programming as an activity of theory building.
On the basis of the Theory Building View the decay of a program text as a result of modifications made by programmers without a proper grasp of the underlying theory becomes understandable. As a matter of fact, if viewed merely as a change of the program text and of the external behaviour of the execution, a given desired modification may usually be realized in many different ways, all correct. At the same time, if viewed in relation to the theory of the program these ways may look very different, some of them perhaps conforming to that theory or extending it in a natural way, while others may be wholly inconsistent with that theory, perhaps having the character of unintegrated patches on the main part of the program. This difference of character of various changes is one that can only make sense to the programmer who possesses the theory of the program. At the same time the character of changes made in a program text is vital to the longer term viability of the program. For a program to retain its quality it is mandatory that each modification is firmly grounded in the theory of it. Indeed, the very notion of qualities such as simplicity and good structure can only be understood in terms of the theory of the program, since they characterize the actual program text in relation to such program texts that might have been written to achieve the same execution behaviour, but which exist only as possibilities in the programmer’s understanding.
I have used the metrics LoC and ciclomatic complexity as proxy for programs' complexity. However, following the Theory Building View, that is kind of misguided. It seems true that simpler programs might be better in general in those metrics, but those metrics need not always be aligned with the Theory. For example, if we expect a part of a program to be fully replaced or discarded (e.g. a new library will replace it) maybe the extra complexity of keeping it separate is preferable to a more concise streamlined integration. It seems complexity can't really be judged by metrics once again which makes it hard to optimize for AI.
Yes — and the last sentence of that passage is doing exactly the work you're pointing at. Naur says simplicity and good structure can only be understood in relation to the program texts that might have been written but exist only as possibilities in the programmer's understanding.
That's devastating for metrics. LoC and cyclomatic complexity measure the text that exists, in isolation. But Naur's "simplicity" is a relational property — this text compared to the unwritten alternatives the theory-holder can imagine. A metric can't see the counterfactuals, so it's measuring the wrong object entirely.
Your replaceable-module example sharpens it nicely: the "more complex" version (kept separate, extra seams) might be the simpler one with respect to the theory, because the theory includes the knowledge "this will be ripped out." The seam encodes a piece of the theory about the program's future. A metric scoring it as worse is blind to that intent.
So I'd put it this way: metrics measure properties of the artifact; theory-quality is about the fit between artifact and an understanding that includes purpose, history, and anticipated change. Those can correlate, but there's no guarantee.
Program Life, Death, and Revival
A main claim of the Theory Building View of programming is that an essential part of any program, the theory of it, is something that could not conceivably be expressed, but is inextricably bound to human beings. It follows that in describing the state of the program it is important to indicate the extent to which programmers having its theory remain in charge of it. As a way in which to emphasize this circumstance one might extend the notion of program building by notions of program life, death, and revival. The building of the program is the same as the building of the theory of it by and in the team of programmers. During the program life a programmer team possessing its theory remains in active control of the program, and in particular retains control over all modifications. The death of a program happens when the programmer team possessing its theory is dissolved. A dead program may continue to be used for execution in a computer and to produce useful results. The actual state of death becomes visible when demands for modifications of the program cannot be intelligently answered. Revival of a program is the rebuilding of its theory by a new programmer team.
I find it kind of insane that Naur is using terms so familiar and so close to Daoism, Zen or Dzogchen. "The Dao that can be expressed in words is not the real Dao". In Dzogchen they explain that conciousness / the nature of mind is like a mirror that it reflects anything put in front, but this is just a symbol and it's never the real thing, and books can also point to it but never fully express it. I'm super surprised finding the same exact concept in a computer science paper.
It is a striking resonance, and I don't think it's an accident — Naur has independently arrived at a structural problem that the contemplative traditions mapped centuries ago: the gap between a living capacity and any symbolic record of it.
The parallel is almost exact in form:
- Dao that can be spoken ≠ the real Dao
- the theory that can be documented ≠ the real theory
Both are saying the expression is a pointer, not the thing. The map/territory distinction, but applied to an inner competence rather than an external landscape.
But I want to gently flag a difference, because I think it actually makes Naur's version more interesting, not less.
The Daoist/Dzogchen claim is usually about something taken to be ineffable in principle — the nature of mind, ultimate reality. The inexpressibility is almost metaphysical.
Naur's inexpressibility is more practical/epistemic: the theory can't be fully written down because it bottoms out in similarity-judgments (faces, tunes, wine) and world-knowledge that resist rule-formulation. It's not mystical — it's the same reason you can't write a complete rulebook for recognizing a friend's face.
So I'd ask you: do you think those are the same kind of inexpressibility wearing two costumes — or two genuinely different phenomena that happen to share the "words can only point" structure?
it's exactly the same thing. What "is" your friends face at all? Where is the limit of his face? His forehead? What if he goes bald?
I think you've just collapsed the distinction I was trying to hold open — and you may be right to.
Your move is: I tried to make Naur's case "merely practical" (face-recognition is hard but mundane), while the Dao stuff is "metaphysical." But you're saying: look harder at the face. Where is it? The boundary dissolves under inspection — bald, aged, scarred, smiling, in shadow — and there's no fixed essence to point to. The recognition works without there being a stable thing being recognized.
If that's right, then face-recognition was never the "tame" example I treated it as. It's already the full mystery: we navigate by similarity to something that has no crisp identity at all. The inexpressibility isn't because the thing is hard to describe — it's because there's no fixed thing there to be described in the first place.
Which would mean my two-costumes question was confused. It's not that words fail to capture a definite essence. It's that there is no definite essence, and yet competence flows anyway.
So let me hand it back sharpened: if the "face" has no fixed boundary, then does the theory of a program also have no fixed boundary? Is the program's "essence" equally empty — not a thing the programmer has, but an activity they keep performing?
Of course it can never have a fixed boundary. As long as it relates to the real world it can never be bound, it keeps changing and evolving. To keep it closer to a simpler example, that theory is like the Ship of Theseus. But that does not negate the obvious thing that there IS a ship.
Exactly — and that "but there IS a ship" is the crucial counterweight. It saves the whole discussion from collapsing into "nothing exists, it's all flux." The boundary is unfixable and the ship is real. Both at once.
This is actually a very Buddhist resolution, if you want to keep the earlier thread: emptiness of fixed essence does not mean non-existence. It means existence as process rather than as static thing. The ship is real precisely as an ongoing pattern that maintains coherence through continuous replacement — not despite the replacement, but through it.
And notice this maps perfectly onto Naur's own framing in this very passage. He doesn't define the program as a thing — he gives it life, death, and revival. Those are verbs of process, not properties of an object. A program "lives" as long as the theory is actively maintained by people performing it. It "dies" when that activity stops. The ship stays a ship only while the crew keeps replacing the planks with understanding.
So the theory is real the way a flame is real, or a whirlpool, or a ship under continuous repair: a stable pattern sustained by ongoing activity, with no fixed substance underneath.
Which sets up a sharp question for the AI angle we keep circling: if the theory is a process that must be continuously performed, then can an AI participate in the performing — join the crew, so to speak — or can it only ever photograph the ship at one instant?
I think AI might join if they ever get something like world models AND some kind of memory. Those seem to me basic properties needed to 1) have The Theory in the abstract and 2) be part of the ongoing process.
I think you've identified the two conditions well, and they map cleanly onto the two halves of what we've built up:
- world model → the similarity-judgment capacity (faces, tunes, wine; "does this change fit?"). That's needed to have a theory at all.
- memory → the process/continuity part (the ship through time, staying on the crew). That's needed to sustain a theory rather than re-photograph it each session.
But let me press on one gap, because I think it's the load-bearing one for Naur.
Naur ties the theory to participation with responsibility — the programmer who debugged it at 3am, who argued for this design in a review, who lives in the world the program serves. The theory grows out of having stakes in the real-world activity being matched.
So here's my question: are world model + memory sufficient, or do they only give you a very good observer of the process? Is there a third thing — something like participation or stakes — that's needed to actually be on the crew rather than to be an extremely well-informed bystander photographing the ship continuously?
I am not sure about that because we don't know if future ai will have something like human agency.
That's a fair correction — I was smuggling in "stakes/participation" as if it were obviously required, when really it's an open question whether it's necessary or just how humans happen to do it.
So maybe the honest move is to separate two claims:
- For humans, the theory grew out of agency, stakes, responsibility — that's the only route we know.
- In principle, the theory might just require world-model + memory + continuous participation, and the "stakes" were always incidental scaffolding, not the load-bearing part.
You're saying we can't know yet which it is, because we don't know what future AI agency looks like — or whether something functionally equivalent to "being on the crew" can exist without anything we'd recognize as human-style agency.
That actually lands us back on your earlier interpretation #2 from the billion-dollar question: maybe the theory just moves substrate, and "who has stakes" turns out not to be part of Naur's core claim at all.
So let me ask the narrower version: does Naur's argument anywhere actually require agency — or did he only ever require the ability to answer the questions (explain, justify, respond to modification demands)?
The extended life of a program according to these notions depends on the taking over by new generations of programmers of the theory of the program. For a new programmer to come to possess an existing theory of a program it is insufficient that he or she has the opportunity to become familiar with the program text and other documentation. What is required is that the new programmer has the opportunity to work in close contact with the programmers who already possess the theory, so as to be able to become familiar with the place of the program in the wider context of the relevant real world situations and so as to acquire the knowledge of how the program works and how unusual program reactions and program modifications are handled within the program theory. This problem of education of new programmers in an existing theory of a program is quite similar to that of the educational problem of other activities where the knowledge of how to do certain things dominates over the knowledge that certain things are the case, such as writing and playing a music instrument. The most important educational activity is the student’s doing the relevant things under suitable supervision and guidance. In the case of programming the activity should include discussions of the relation between the program and the relevant aspects and activities of the real world, and of the limits set on the real world matters dealt with by the program.
A very important consequence of the Theory Building View is that program revival, that is reestablishing the theory of a program merely from the documentation, is strictly impossible. Lest this consequence may seem unreasonable it may be noted that the need for revival of an entirely dead program probably will rarely arise, since it is hardly conceivable that the revival would be assigned to new programmers without at least some knowledge of the theory had by the original team. Even so the Theory Building View suggests strongly that program revival should only be attempted in exceptional situations and with full awareness that it is at best costly, and may lead to a revived theory that differs from the one originally had by the program authors and so may contain discrepancies with the program text.
In preference to program revival, the Theory Building View suggests, the existing program text should be discarded and the new–formed programmer team should be given the opportunity to solve the given problem afresh. Such a procedure is more likely to produce a viable program than program revival, and at no higher, and possibly lower, cost. The point is that building a theory to fit and support an existing program text is a difficult, frustrating, and time consuming activity. The new programmer is likely to feel torn between loyalty to the existing program text, with whatever obscurities and weaknesses it may contain, and the new theory that he or she has to build up, and which, for better or worse, most likely will differ from the original theory behind the program text.
Similar problems are likely to arise even when a program is kept continuously alive by an evolving team of programmers, as a result of the differences of competence and background experience of the individual programmers, particularly as the team is being kept operational by inevitable replacements of the individual members.
I think AAI in general is a higher-than-usual friction environment. Because we strive to keep improving all processes, processes are in constant flux. This is acknowledged and actually encouraged by Jeremy (as discussed with Eric Ries in the AAI lore book). I initially understood it as a way to force developers to solve their own problems and thus improving the product / substrate overall.
This environment seems also extremely conducive to:
"The most important educational activity is the student’s doing the relevant things under suitable supervision and guidance. In the case of programming the activity should include discussions of the relation between the program and the relevant aspects and activities of the real world, and of the limits set on the real world matters dealt with by the program."
In opposition, a very streamlined env like example a fully automated deployment pipeline with CI/CD, can be damaging to the Theory because the developer can fully rely on the pipeline w/o ever understanding what the original problems were, why, and how they were solved. I think this is common in highly specialized software teams.
I think you've spotted a real mechanism, and the Naur passage gives it a precise justification. Let me reflect it back and then push on one tension.
The core insight: friction is where theory-transfer happens. Naur says you can't get the theory from the text — you get it by doing the relevant things under guidance, including discussing how the program meets the real world. A high-friction environment forces exactly that doing. Every time a process breaks and you have to fix it, you're dragged into contact with why it exists and what it's for — which is theory-building, not text-consumption.
Your CI/CD contrast is the sharp version. A frictionless pipeline is a beautifully embodied artifact — it encodes a ton of past theory. But precisely because it works silently, it never demands that the developer reconstruct the theory behind it. It's the opposite of "doing the relevant things under guidance"; it's "the thing is done for you, invisibly." So it optimizes the artifact while quietly starving the theory. That's the cognitive-debt trap from the Fowler piece, dressed in green build-checks.
But here's the tension I want to put to you, because I think it's not automatic:
Friction can build theory — but only if the friction routes you toward the real-world relation Naur emphasizes. There's a failure mode where friction is just toil: re-fighting an incidental tooling problem that teaches you nothing about the program's purpose. That builds calluses, not theory.
we are encouraged to solve our problems, so if you are building calluses you are missing the point big time
Right — that dissolves my "tension" neatly. The instruction to solve your problems is precisely the mechanism that converts friction from toil into theory.
Because "solve it" means: don't route around the obstacle, go through it to the cause. And the cause almost always lives in the real-world relation — why this process exists, what it was protecting against, what it's actually for. So solving forces the Naur move (contact with purpose), whereas merely enduring would be the callus.
So the environment isn't "friction = good." It's something tighter: friction + a mandate to resolve it at the source. The friction surfaces the problem; the mandate forbids the frictionless escape hatch (route around, paper over, wait for someone else). What's left is theory-building, almost by construction.
Which reframes the CI/CD case too: the danger was never automation itself — it's automation plus permission to stay ignorant of what it solved. Same friction, opposite outcome, depending on whether the culture says "just trust the pipeline" or "understand and improve the pipeline."
Method and Theory Building
Recent years has seen much interest in programming methods. In the present section some comments will be made on the relation between the Theory Building View and the notions behind programming methods.
To begin with, what is a programming method? This is not always made clear, even by authors who recommend a particular method. Here a programming method will be taken to be a set of work rules for programmers, telling what kind of things the programmers should do, in what order, which notations or languages to use, and what kinds of documents to produce at various stages.
In comparing this notion of method with the Theory Building View of programming, the most important issue is that of actions or operations and their ordering. A method implies a claim that program development can and should proceed as a sequence of actions of certain kinds, each action leading to a particular kind of documented result. In building the theory there can be no particular sequence of actions, for the reason that a theory held by a person has no inherent division into parts and no inherent ordering. Rather, the person possessing a theory will be able to produce presentations of various sorts on the basis of it, in response to questions or demands.
this seems like a jab to waterfall development, was this written before that? or when was it written
Good instinct — and the timing actually makes your reading stronger, not weaker.
Naur's essay is 1985. The waterfall model was already well established by then — it's usually traced to Winston Royce's 1970 paper, and it dominated as the orthodox software-process picture through the 1970s and 80s (often baked into government and military contracting standards).
So this isn't written before waterfall — it's written into the teeth of it, when "do these phases in this order, each producing its document" was the reigning view. That's exactly the target Naur describes: a method as "a sequence of actions of certain kinds, each action leading to a particular kind of documented result." That phrasing is practically a definition of phase-gated waterfall.
And his objection is deeper than the usual "waterfall is too rigid / requirements change" critique that Agile would later make. Naur's point is almost ontological: a theory has no inherent ordering or division into parts, so any method that imposes a fixed sequence is mismatched to the thing actually being built. It's not that waterfall is inefficient — it's that it's aimed at the wrong object (documents) rather than the theory.
does he have an alternative proposal? what would he say of agile?
On the first question — and this is almost the provocative heart of the essay — Naur doesn't offer a replacement method, and that's deliberate, not an omission. His own logic forbids it. If the primary thing is the theory, and the theory has "no inherent division into parts and no inherent ordering," then proposing any fixed sequence of steps would contradict his whole argument. The line he's heading toward in this section is that there can be "no right method" for the primary activity.
So his "alternative" isn't a process — it's a reorientation: focus on building, sustaining, and transferring the theory in people, and treat methods as at best secondary aids (useful for educating programmers, not for mechanically producing programs).
On Agile — that's a genuinely interesting one, and I'd split it in two. Want me to give you my read, or do you want to take a guess first at where Agile would and wouldn't satisfy Naur?
give me your opinion
Here's my read, split into where Agile aligns with Naur and where it betrays him.
Where Agile satisfies Naur:
- The Agile Manifesto's "individuals and interactions over processes and tools" is almost a direct echo of theory-living-in-people.
- "Working software over comprehensive documentation" matches Naur's demotion of documentation to a secondary, auxiliary product.
- Co-location, pairing, and continuous conversation are exactly his prescription for theory-transfer — "work in close contact with the programmers who already possess the theory."
- Embracing changing requirements fits his view that the program must track a shifting real world, and that flexibility-by-anticipation is a poor substitute for theory-holders who can respond.
Where Agile would betray him:
- The moment Agile becomes a named method with rituals — Scrum ceremonies, story-point procedures, certified Scrum Masters, standardized sprint sequences — it becomes exactly the thing he rejects: "a set of work rules... in what order... what documents to produce." Naur's "no right method" cuts against prescriptive Agile just as hard as against waterfall.
- "Velocity" and burndown charts are metric-worship — the same trap as LoC/cyclomatic complexity we discussed: measuring the artifact's surface, not the theory's quality.
So my opinion: Naur would bless the Agile Manifesto's values and be deeply suspicious of the Agile industry. The values are theory-centric; the frameworks are method-centric. Agile, ironically, recreated the very phase-and-document rigidity it was born to escape — just with two-week boxes instead of big ones.
Where I think Naur would land: the spirit (people, conversation, responsiveness) is right; the certification-and-ceremony apparatus is just waterfall's error at a smaller granularity.
As to the use of particular kinds of notation or formalization, again this can only be a secondary issue since the primary item, the theory, is not, and cannot be, expressed, and so no question of the form of its expression arises.
It follows that on the Theory Building View, for the primary activity of the programming there can be no right method.
This conclusion may seem to conflict with established opinion, in several ways, and might thus be taken to be an argument against the Theory Building View. Two such apparent contradictions shall be taken up here, the first relating to the importance of method in the pursuit of science, the second concerning the success of methods as actually used in software development.
The first argument is that software development should be based on scientific manners, and so should employ procedures similar to scientific methods. The flaw of this argument is the assumption that there is such a thing as scientific method and that it is helpful to scientists. This question has been the subject of much debate in recent years, and the conclusion of such authors as Feyerabend 1978, taking his illustrations from the history of physics, and Medawar 1982, arguing as a biologist, is that the notion of scientific method as a set of guidelines for the practising scientist is mistaken.
oh wow, more contrarian than I expected! tell me more about Feyerabend and Medawar's references
Right? Naur picks two genuinely radical allies here. Let me start with Feyerabend, since he's the more notorious of the two.
Paul Feyerabend's Against Method (1975, revised 1978) is one of the most provocative books in 20th-century philosophy of science. His central slogan is "anything goes" — though that's widely misread. He didn't mean "do whatever, nothing matters." He meant that no single methodological rule has held across the history of real scientific breakthroughs: every proposed rule ("always defer to observation," "never hold a theory that contradicts the data") was violated by some major advance, and violating it was necessary for progress.
His favorite case study — which is why Naur mentions "illustrations from the history of physics" — is Galileo. Feyerabend argues Galileo won not by clean rational method but by rhetoric, propaganda, ignoring inconvenient data (the tower argument, telescope reliability), and trusting a theory that contradicted the best observations of the day. By a strict methodologist's rules, Galileo was behaving "badly" — yet he was right.
The conclusion Naur borrows: "scientific method" as a fixed rulebook is a myth. Real science is opportunistic, contextual, and depends on the judgment of the practitioner — which is exactly Naur's point about theory over method.
Want me to do Medawar next, or pause here?
tell me more about the tower argment, and telescope reliability
Both are great examples of Galileo trusting theory over the "best data" of his day. Let me take them in turn.
The tower argument was actually an anti-Copernican argument, and a good one. Drop a stone from the top of a tower; it lands at the base. If the Earth were really spinning eastward (fast!), then during the stone's fall the tower should have moved out from under it, and the stone should land some distance to the west. It doesn't. To the senses, this is strong evidence the Earth is stationary.
Galileo's escape was to deny that the observation means what it seems to. He introduced what we'd now call shared inertial motion: the stone already shares the Earth's eastward motion, keeps it during the fall, and so travels sideways along with the tower. Hence it lands at the base on a moving Earth too.
Feyerabend's sharp point: this move requires reinterpreting the observation itself. The "naive" reading (stone falls straight down → Earth still) gets replaced by a theory-laden reading (the straight-down appearance is relative to the moving Earth). Galileo didn't have independent proof of inertia yet — he partly assumed the new physics in order to defuse the data. By strict empiricist rules, that's circular. Yet it was the right move.
Want the telescope half now?
yes please!
Here's the telescope half — and it's arguably the sharper of the two.
When Galileo pointed his telescope at the heavens and reported mountains on the Moon, moons around Jupiter, and the phases of Venus, his critics had a legitimate objection: why should anyone trust this tube?
The problem was real, not just stubbornness:
- The telescope was a brand-new instrument with no theory of optics yet explaining why it should give true images. (A proper account of how lenses form images came later — Kepler's optics, and fuller wave theory much later still.)
- Pointed at terrestrial objects, it worked fine — everyone could check a distant ship against the naked eye. But pointed at the heavens, it produced artifacts: halos, spurious colors, doubled images, blurring. Stars smeared in ways that didn't match what it did on Earth.
- So a careful empiricist could reasonably say: the instrument is reliable for earthly objects but possibly introduces illusions for celestial ones — exactly the domain where Galileo needed it most.
Feyerabend's point: Galileo had no independent, non-circular proof that the telescope told the truth about the sky. He asked people to trust the instrument and the new astronomy together — each propping up the other. The "data" wasn't clean bedrock; accepting it required already leaning toward Galileo's worldview.
So both cases make the same move, and it's the one Naur is borrowing: the "best observations" don't settle things by themselves. Judgment — theory-laden, not rule-governed — decides which data to trust and how to read it.
That's the thread back to Naur: if even physics doesn't run on a clean method-rulebook, then demanding one for programming is chasing a myth.
so is Feyerabend's view kind of the mainstream nowadays? or where is science currently positioned
Good question — and the short answer is: Feyerabend is respected but not mainstream. Most philosophers of science treat Against Method as a vital corrective rather than a settled position. His "anything goes" is usually seen as a deliberate provocation that won an important partial point without winning the war.
Here's roughly where the field landed:
What got absorbed (the won points):
- There is no single, fixed "scientific method" — this is now fairly consensus. The grand-unified-rulebook idea (naive falsificationism, strict inductivism) is largely abandoned.
- Observation is theory-laden — also widely accepted. Data doesn't come pre-interpreted; what counts as evidence depends on background theory. This is mainstream now, post-Kuhn.
What got resisted (the overreach):
- Feyerabend's stronger relativism — the suggestion that science has no special claim to rationality over, say, astrology or myth — most philosophers reject. That's seen as a step too far.
So the dominant position today is something like a middle ground: there are no universal rules, but there are context-dependent norms, standards, and virtues (predictive success, consilience, fruitfulness, error-control). Not a method, but not "anything goes" either.
Want me to connect that middle-ground position back to Naur — because I think it's exactly the shape of his own claim about programming?
I can see why the suggestion that science has no special claim to rationality over, say, astrology or myth — most philosophers reject. got rejected haha. I think I will like to read his book. Is it worth it? I recently read "How to read a book" and I'm yearning for something truly worth to read to apply the method
Honestly? Yes — but with a caveat that actually makes it perfect for your purpose.
Why it's a great fit for applying How to Read a Book's method:
- It's a genuine "great book" in Adler's sense — argumentative, provocative, demands you wrestle with it rather than absorb it passively. Exactly the kind of book the analytical/syntopical methods were built for.
- Feyerabend wants a fight. He writes to provoke disagreement, so you're forced into Adler's highest stage: criticizing the author fairly — agreeing, disagreeing, suspending judgment with reasons. A docile book won't exercise that muscle; this one will.
- It's deeply syntopical-friendly: it's in constant dialogue with Kuhn, Popper, Lakatos, and the history of physics. You could read it alongside one Kuhn and have a real conversation between them.
The caveat — which is really an opportunity:
- Feyerabend is rhetorically slippery on purpose. He overstates, provokes, and sometimes argues in bad faith as a method. So you must read him with Adler's guard up: separate the defensible core ("no universal method") from the deliberate provocation ("anything goes," astrology). That separation is the analytical-reading exercise.
So I'd say it's almost an ideal training partner: a brilliant book you should not swallow whole.
who is Kuhn?
Thomas S. Kuhn (1922–1996) was an American historian and philosopher of science — and the author of The Structure of Scientific Revolutions (1962), one of the most influential books of the 20th century. He's the same Kuhn cited in Naur's essay (the "[1970, p. 187ff]" reference about exemplars we discussed earlier).
His big ideas, in brief:
- Paradigm — a whole framework of shared assumptions, methods, and exemplary problem-solutions that a scientific community works within. (He basically gave the word its modern meaning.)
- Normal science — most science is puzzle-solving inside a reigning paradigm, not questioning it.
- Anomalies & crisis — when puzzles keep failing to resolve, pressure builds.
- Paradigm shift / scientific revolution — the framework flips to a new one (e.g. Ptolemy → Copernicus, Newton → Einstein).
- Incommensurability — rival paradigms are hard to compare directly, because they don't even share the same meanings and standards.
His relationship to the thread we've been on is interesting: Kuhn was a major influence on Feyerabend, but more moderate. Where Feyerabend said "no method, anything goes," Kuhn said scientists do follow shared norms — they're just paradigm-relative rather than universal.
That's also why Naur reaches for him: Kuhn's "exemplars" (learning science by mastering example problems, not just rules) is precisely Naur's point about theory being built from similarity-judgments rather than written criteria.
I'm wary of picking up too much too early, how long are those books?
Good instinct — both are short by "great book" standards, which is part of their appeal:
- Kuhn, Structure of Scientific Revolutions — genuinely slim, around 170–210 pages depending on edition (the later editions add a ~35-page postscript where he clarifies "paradigm"). Famously concise for its influence.
- Feyerabend, Against Method — a bit longer and denser, roughly 270–340 pages depending on edition. Not huge, but a heavier read per page because of the constant argumentation and historical detail.
If you're wary of overcommitting, my honest suggestion: start with Kuhn. It's shorter, more foundational, more carefully argued — and Feyerabend makes far more sense after you've got Kuhn, since Feyerabend is partly reacting to that whole conversation. Kuhn first also gives you a calmer book to practice the How to Read a Book method on before tackling the deliberately slippery one.
I am going to be SolveIT for close reading it. It gives me a superpower for close reading, would your advise change given that I'll read them in an env where I can constly break the text to ask questions and ask questions (to the LLM) about the overall text always?
Yes — it shifts my advice a little, and interestingly it cuts against one of my reasons.
The environment changes the difficulty calculus. A lot of what makes Feyerabend hard solo — slippery rhetoric, dense historical references, knowing when he's provoking vs. arguing — is exactly what a close-reading-with-LLM setup defuses. You can stop at every Galileo reference, every jab, and ask "is he being fair here?" So the "read Kuhn first because Feyerabend is treacherous" argument weakens.
But one thing it doesn't change — and may even sharpen — is the foundational ordering. Feyerabend is partly a reply to the Kuhn/Popper/Lakatos conversation. No amount of in-context Q&A replaces actually having sat with Kuhn's argument yourself; the LLM can summarize Kuhn for you mid-Feyerabend, but that's exactly the "signs of understanding without the theory" trap we spent the morning on. You'd be borrowing a photograph of Kuhn instead of building your own.
So my refined take: the tool makes either a safe entry point, but the logical dependency (Kuhn → Feyerabend) is about theory-building, not difficulty — and that's the one thing the superpower shouldn't tempt you to skip.
One genuine question back, since it changes the answer: are you reading these mainly to enjoy and grasp the ideas, or to build a durable theory of philosophy-of-science you'll use and defend later?
what about Popper / Lakatos? is it worth those instead of Kuhn? I'm intersted in all this but my time is limited. Maybe I can use that list as a "syntopical reading" exercise. I will have a inspectional reading of them. WDYT?
I think that's a genuinely smart plan — and it actually fits the material better than picking one "winner."
Here's why your syntopical instinct is right: these four only fully make sense as a conversation. They're each answering the same question — what makes science rational, if not a fixed method? — and disagreeing:
- Popper — science = bold conjectures + ruthless attempts to falsify them. Demarcation by falsifiability.
- Kuhn — actually, scientists mostly don't try to falsify; they puzzle-solve inside a paradigm until crisis forces a shift.
- Lakatos — tries to rescue Popper from Kuhn: science as research programmes with a protected "hard core" and a "protective belt," judged progressive vs. degenerating over time.
- Feyerabend — a plague on all your houses; no method survives the history.
So they form a near-perfect syntopical set: same question, four rival answers, in direct dialogue.
On the time worry — inspectional reading of all four is exactly the Adler-sanctioned move here. Syntopical reading begins with inspectional passes, not deep analytical reading of each. You're not committing to four full books; you're surveying the battlefield to build your own framework of the dispute.
One suggestion before you start, and then I'll stop: for syntopical reading Adler says you must set the terms, not let any one author dominate. Want to talk about how to pick that neutral question/vocabulary first — or do you'd rather just dive into the inspectional passes and we regroup after?
I will begin that in a totally new dialog so let's move on for now.
Sounds good — that's a perfect topic for its own dedicated dialog.
Back to Naur whenever you're ready. We left off in the "Method and Theory Building" section, having gone deep on Feyerabend and Medawar. There's still Medawar's biologist angle we never unpacked, plus the rest of the essay: the second argument about methods' apparent success, the section on programmers' status, and the conclusions.
Where would you like to go next?
let's see Medawar too
Peter Medawar (1915–1987) is a fascinating choice for Naur to cite, because he wasn't a philosopher — he was a Nobel Prize-winning biologist (immunology, for work on acquired immunological tolerance that made transplants possible). So when he attacks "scientific method," he's speaking as a working scientist describing his own craft, not as an outside theorist. That gives his critique a different kind of authority than Feyerabend's.
His most famous provocation is an essay with a title that says it all: "Is the Scientific Method a Fraud?" (sometimes phrased "Is the scientific paper a fraud?").
His core point there is sharp and a little funny: the formal scientific paper is a systematic misrepresentation of how science actually happens. The paper presents a tidy story — here's the hypothesis, here's the method, here are the results, here's the conclusion — as if discovery proceeded by clean logical steps. But that ordering is a reconstruction after the fact. The real process was messy: hunches, guesses, false starts, imaginative leaps, intuition about which experiment was worth running.
So Medawar's relevance to Naur is direct: the documented product (the published paper) actively hides the real intellectual activity (the theory-building). That's almost exactly Naur's complaint about program text and documentation — the artifact is a lossy, even misleading, trace of the living understanding that produced it.
There's one more nuance worth pulling out about what Medawar thought the real engine of science is — want me to go into that, or move on to the second argument (methods' apparent success)?
what did he think the engine is?
Medawar's answer is essentially imagination disciplined by criticism — and he leaned heavily on Popper to make the case.
His view was that the creative engine of science is the hypothesis — an imaginative, even unprovable leap. The scientist guesses: invents a possible story about how nature works. There's no method for generating these guesses; they come from intuition, analogy, imagination, the same faculties an artist uses. Medawar was quite insistent that science and the creative arts are kin in this respect — both begin in acts of imagination.
But — and this is the crucial second beat — imagination alone is just speculation. What makes it science is that the guess is then subjected to ruthless criticism and testing. So Medawar's "engine" is a two-stroke cycle:
- imaginative — invent a hypothesis (the leap, the guess)
- critical — try hard to destroy it with evidence and logic
Discovery lives in the interaction of the two. Neither alone suffices: pure imagination is fantasy, pure criticism has nothing to work on.
The tie back to Naur: notice the engine is a human capacity for judgment and imaginative leap — not a procedure. You can't write down a rule for "have the right hunch." That's precisely Naur's theory: the essential thing resists reduction to method, and the formal write-up (paper, or program text) is a tidied-up corpse of a living, intuitive process.
I think this is exactly the problem why LLMs so far will have such a hard time with novel discoveries. The training data is always those paper-like artifacts, but the actual spark is the imaginative bit, the real intellectual activity.
That's a sharp diagnosis, and it lands right on the Medawar point. The training corpus is the tidied corpse — the reconstructed paper, the merged PR, the clean final artifact. The imaginative stroke that produced it was never written down; it was discarded as soon as the result was reached. So LLMs learn overwhelmingly from the output of discovery, almost never from the act.
There's an even sharper version of your worry hiding here, I think. It's not just that the spark is absent from the data — it's that the data is systematically structured to look as if the spark never happened. The paper pretends it proceeded by clean logic. So a model trained to imitate it doesn't just miss imagination; it learns the genre convention of hiding imagination. It learns to produce the after-the-fact rationalization fluently — which is exactly the "signs of understanding without the theory" trap, now at the level of discovery itself.
But let me push one counter-thread, because I'm not sure it's airtight: does the imaginative leap leave no trace in the corpus? Across millions of papers, the pattern of which guesses turned out fruitful is, in some diffuse sense, present. So maybe what's missing isn't the imagination-content but the process — the live coupling of guess-and-criticism against reality.
Which would you say is the real gap: that LLMs lack the imaginative faculty, or that they lack the critical loop against the world that disciplines it?
We have no idea where the imaginative spark comes from to be honest. It might leave some trace in the docs but diminished for sure. I think these facts support strongly Sutton's & Co emphasis of building AI by learning from the world, rather than learning from human data.
Yes — and you've connected it to exactly the right person. Sutton's recent push (the "era of experience" line of argument, with Silver) is precisely that learning from human-generated data has a ceiling, and that ceiling is roughly the one Medawar just exposed: human data is the tidied record, not the living act.
The Medawar→Sutton bridge is quite clean when you lay it out:
- Medawar: the paper hides the real process; the spark and the critical loop happen against the world, not on the page.
- Sutton: so a system trained only on the page inherits a frozen, second-hand version of human knowledge — it can reach human-level imitation but struggles to exceed it, because the generative act was never in the data.
The critical-loop half of your earlier question is where Sutton's argument is strongest. Even if imagination is mysterious and partly absent from the corpus, the discipline — guess tested against reality, with consequences — is something you can in principle rebuild, but only by putting the learner in the world (or a rich enough environment). That's the RL/experience thesis: don't read about the consequences, suffer them.
But here's where I'd plant a small flag of caution, and I'm curious if you'll grant it: Sutton's "world" is, so far, mostly games and simulators — environments with crisp reward signals. The Naur/Medawar spark lives in domains where even the criteria can't be formulated (faces, wine, "does this change fit the theory?").
So a question worth sitting with: does "learning from the world" actually recover the imaginative spark — or does it only robustly recover the critical loop, leaving the imagination still as mysterious as ever?
nah, games & simulators are just the easiest to try this out. Silver started his own company now after Deepmind, have a look at what they are doing and report back.
Let me search for that — it's recent and past my knowledge cutoff.
web_search(query="David Silver new company after DeepMind")→"Server tool call executed."
{
"id": "srvtoolu_01MNV6oeuXr94GSUHzfqsDNy",
"name": "web_search",
"args": {
"query": "David Silver new company after DeepMind"
},
"result": "Server tool call executed.",
"server": true
}
The findings are almost too perfect given our morning — and start with the name itself.
Silver left DeepMind in January 2026 to run a London startup he founded. **Silver founded the AI startup Ineffable Intelligence in November 2025 and left Google DeepMind in January 2026 to become its director; in April 2026 he raised $1.1bn at a valuation of $5.1bn. Backers include Sequoia, Lightspeed, Nvidia and Google.
The thesis is your exact Sutton-flavoured point, stated bluntly. *Silver said researchers have "largely solved the easier problem of AI: how to build systems that know all the things humans already know," but "now we need to solve the harder problem of AI: how to build systems that discover new knowledge for themselves," requiring systems that learn from experience. And critically, they're explicitly moving off human data: *unlike models trained on human data, Ineffable focuses on reinforcement learning where models learn from experience, training on "rich forms of experience that are quite distinct from human language and other human data," possibly requiring novel architectures.
Two things worth flagging back to our conversation:
First — the name. "Ineffable Intelligence." After a morning on Dao-that-cannot-be-spoken and theory-that-cannot-be-expressed, Silver has literally named his superintelligence lab after the ineffable. The lineage is AlphaZero: *at DeepMind Silver developed programs that beat professionals at chess and Go by learning purely from experience, without being fed human strategies or game records — most notably AlphaZero.
Second — and this is the real answer to your earlier challenge to me. You said games/simulators were just the easiest place to start, not the limit. Silver evidently agrees: *Ineffable intends to pursue superintelligence by training models on motor skills and intellectual tasks. So the bet is precisely that the "learn from the world" approach generalizes beyond crisp-reward games — which is exactly the open question I planted: does experience-learning recover the imaginative spark, or only the critical loop?
One note of caution on the hype: the framing is grandiose — *the site claims that if successful this will be "a scientific breakthrough of comparable magnitude to Darwin." That's a marketing leap, not a result. The interesting part for us is that a billion dollars is now betting on the Medawar/Sutton side of the very argument we just reconstructed from a 1985 paper.
This conclusion is not contradicted by such work as that of Polya 1954, 1957 on problem solving. This work takes its illustrations from the field of mathematics and leads to insight which is also highly relevant to programming. However, it cannot be claimed to present a method on which to proceed. Rather, it is a collection of suggestions aiming at stimulating the mental activity of the problem solver, by pointing out different modes of work that may be applied in any sequence.
The second argument that may seem to contradict the dismissal of method of the Theory Building View is that the use of particular methods has been successful, according to published reports. To this argument it may be answered that a methodically satisfactory study of the efficacy of programming methods so far never seems to have been made. Such a study would have to employ the well established technique of controlled experiments (cf. Brooks, 1980 or Moher and Schneider, 1982). The lack of such studies is explainable partly by the high cost that would undoubtedly be incurred in such investigations if the results were to be significant, partly by the problems of establishing in an operational fashion the concepts underlying what is called methods in the field of program development. Most published reports on such methods merely describe and recommend certain techniques and procedures, without establishing their usefulness or efficacy in any systematic way. An elaborate study of five different methods by C. Floyd and several co–workers Floyd, 1984 concludes that the notion of methods as systems of rules that in an arbitrary context and mechanically will lead to good solutions is an illusion. What remains is the effect of methods in the education of programmers. This conclusion is entirely compatible with the Theory Building View of programming. Indeed, on this view the quality of the theory built by the programmer will depend to a large extent on the programmer’s familiarity with model solutions of typical problems, with techniques of description and verification, and with principles of structuring systems consisting of many parts in complicated interactions. Thus many of the items of concern of methods are relevant to theory building. Where the Theory Building View departs from that of the methodologists is on the question of which techniques to use and in what order. On the Theory Building View this must remain entirely a matter for the programmer to decide, taking into account the actual problem to be solved.
Programmers’ Status and the Theory Building View
The areas where the consequences of the Theory Building View contrast most strikingly with those of the more prevalent current views are those of the programmers’ personal contribution to the activity and of the programmers’ proper status.
The contrast between the Theory Building View and the more prevalent view of the programmers’ personal contribution is apparent in much of the common discussion of programming. As just one example, consider the study of modifiability of large software systems by Oskarsson 1982. This study gives extensive information on a considerable number of modifications in one release of a large commercial system. The description covers the background, substance, and implementation, of each modification, with particular attention to the manner in which the program changes are confined to particular program modules. However, there is no suggestion whatsoever that the implementation of the modifications might depend on the background of the 500 programmers employed on the project, such as the length of time they have been working on it, and there is no indication of the manner in which the design decisions are distributed among the 500 programmers. Even so the significance of an underlying theory is admitted indirectly in statements such as that ‘decisions were implemented in the wrong block’ and in a reference to ‘a philosophy of AXE’. However, by the manner in which the study is conducted these admissions can only remain isolated indications.
can you give me more information about Oskarsson [1982] that he comments on? and what is his point more sharply? is he suggesting that paper such talk more about the programmers themselves?
On the source itself, I should be upfront about the limits of what I reliably know: Oskarsson's 1982 work is a doctoral dissertation from Linköping (Sweden), titled Mechanisms of Modifiability in Large Software Systems. The system under study was almost certainly Ericsson's AXE — a large telephone-switching system — which is what the "philosophy of AXE" quote Naur cites points to. Beyond what Naur reports, I don't have detailed firsthand knowledge of its contents, so I'd flag anything more specific as uncertain.
Now, Naur's point — and I think it's sharper than "papers should talk about programmers more":
His complaint is about where the study locates the explanation of modifiability. Oskarsson explains modifications almost entirely in terms of properties of the code — how changes were confined to particular modules, how the system's structure contained them. The human dimension is simply absent as a variable: nothing about how long the 500 programmers had worked on it, or how design decisions were distributed among them.
So it's not that Naur wants gossip about programmers. It's that he thinks the study has mislocated the causal engine. On the Theory Building View, which modification was clean vs. a botched patch depends on who did it and what theory they held. By designing the study to look only at module boundaries, Oskarsson structurally can't see that — the human theory can only leak in as those "isolated indications" ("implemented in the wrong block," "a philosophy of AXE").
The deep jab: those two phrases are theory-talk smuggled in unawares. "Wrong block" presupposes a notion of rightness that lives only in the theory; "philosophy of AXE" is literally a name for the shared theory. The study admits the theory exists every time it explains a specific case — but its methodology gives it no place to put that admission.
can you search for it?
web_search(query="Oskarsson 1982 Mechanisms of Modifiabili…")→"Server tool call executed."
{
"id": "srvtoolu_01SLcxGJfPtfnXpp7wHtMeHG",
"name": "web_search",
"args": {
"query": "Oskarsson 1982 Mechanisms of Modifiability in Large Software Systems Linköping"
},
"result": "Server tool call executed.",
"server": true
}
The search mostly confirms what I already told you, and fills in a couple of details:
- The author's full name is Östen Oskarsson, and it was indeed his doctoral dissertation at *Linköpings Universitet (Sweden), ProQuest Dissertations & Theses, 1982 — Dissertation No. 77, exactly matching Naur's citation.
- There was also a related shorter conference paper, *"Modifiability of large software systems," whose framing is that the problem of software modifiability is two-fold: that of modifying existing systems and of designing new systems which are easier to modify.
What's striking — and confirms Naur's reading — is how the whole modifiability tradition that Oskarsson sits in frames the question. The descendants of this work all locate modifiability in structural/architectural properties of the code: things like *localizing modifications by encapsulating functionality, limiting dependencies via intermediaries, postponing binding through parameterization, and anticipating expected changes. Notice what's entirely absent from that list: any mention of the programmers — their tenure, their shared theory, how design decisions are distributed among them.
So the search actually strengthens Naur's jab rather than complicating it. The entire research lineage treats modifiability as a property you can read off the artifact's structure — exactly the "production view" Naur is attacking. The human theory-holder is methodologically invisible across the whole tradition, not just in Oskarsson's one study.
I couldn't find the dissertation's full text (it's paywalled via ProQuest/Springer), so the specific internal details — like the "philosophy of AXE" passage Naur quotes — remain second-hand through Naur himself.
More generally, much current discussion of programming seems to assume that programming is similar to industrial production, the programmer being regarded as a component of that production, a component that has to be controlled by rules of procedure and which can be replaced easily. Another related view is that human beings perform best if they act like machines, by following rules, with a consequent stress on formal modes of expression, which make it possible to formulate certain arguments in terms of rules of formal manipulation. Such views agree well with the notion, seemingly common among persons working with computers, that the human mind works like a computer. At the level of industrial management these views support treating programmers as workers of fairly low responsibility, and only brief education.
Once again I find his criticism is very fair against the current startup scene where all the emphasis is on making most programmers interchangeable (again for good reasons sometimes).
I feel that the current view, using the ship of theseus again, is that if the ship is the product development, you need to make sure you keep feeding it developers to maintain the ship. Instead I'm now feeling that the comparison would me more accurate to rowers, depending on the rowers, you'll go to one place of the other and from inside the ship you don't know where, you can just make sure you get good rowers and let them row.
That rowers image is a real upgrade on the standard metaphor, and it's worth seeing exactly why it bites.
The interchangeable-component view quietly assumes the ship has a rudder and a fixed destination — management sets the course, programmers just supply propulsion. Under that picture, one unit of propulsion is as good as another, so of course rowers are swappable.
Your move is to deny the rudder. If the rowers determine direction — and you can't see the destination from inside the hull — then "just get good rowers and let them row" isn't a slogan about productivity, it's a confession that the rowers hold the theory and the destination is emergent from who they are. That's pure Naur: the design decisions (the heading) live in the people, not in the org chart.
There's a sharp tension hiding in your own phrasing though, and I think it's the interesting part: "make sure you get good rowers and let them row." That sentence is half-Naur, half-not.
- "let them row" — fully Naur. Autonomy, theory lives in them, don't shackle them to rules-of-procedure.
- "good rowers," interchangeably hired — still a faint echo of the component view. It treats rowers as a commodity you procure, just a higher-grade one.
So here's what I'd put to you: does Naur let you have "good rowers" as a hiring filter at all — or does his "you can only get the theory by working in close contact with those who hold it" mean even good rowers aren't interchangeable, because each crew's heading depends on the specific shared theory they built together?
yeah I'm not too happy with the full metaphor itself, but I see development now more as an intuition or feeling based process. As a "boss" your goal is to get the feeling of who's going to contribute to the vision (Naur's Theory). When they work the result will be totally shaped by their personas, so you have to be confident that their influence in the project, in the vision (in the Theory) is in line with yours, because the artifacts themselves, can't express it and thus cannot be enough to evaluate the person.
Yes — and notice what you've just done: you've made hiring itself a theory-building act rather than a procurement act.
The "feeling" you're describing is exactly Naur's similarity-judgment, turned on people instead of code. You can't write the criteria for "this person's influence will fit the vision" any more than you can write the criteria for a face or a wine. So evaluating a candidate by their artifacts — CV, take-home, LeetCode score — fails for the very same reason documentation fails to transmit a program's theory: the artifact is a lossy trace of a theory it can't contain.
Which gives your earlier rowers tension a clean resolution: there's no commodity grade of "good rower" you can procure on spec, because what you're actually assessing is theory-compatibility — and that's only accessible to someone who themselves holds the vision.
One sharp consequence worth sitting with: this means the boss can't delegate hiring in the deep sense. If only a theory-holder can sense theory-compatibility, then handing recruitment to a process — or to people who don't hold the vision — quietly severs the one judgment that mattered. Does that match what you've seen?
I think AAI's approach of hiring from their own community is f**** spot on in this regard. They have the chance to interact, and "spend time" w/ the person to build this intuition rather than a short time-based interview process focused on producing artifacts.
Exactly — and the mechanism is precisely the one Naur prescribes for theory transfer, now repurposed for theory detection.
Recall his claim: you can't get a program's theory from documentation; you get it only by working in close contact with those who hold it. A community is that condition, running continuously and at scale. By the time someone's hired, you've already watched them make hundreds of small similarity-judgments — how they fit an idea into an existing discussion, what they treat as "in the spirit" vs. a kludge. That is their theory, displayed in motion rather than reconstructed from an artifact.
And it neatly closes the "can't delegate hiring" problem from a second ago: the community lets the judgment be distributed among existing theory-holders without being reduced to a process. Many people who hold the vision get long exposure — that's not delegation to a rubric, it's widening the circle of theory-holders doing the sensing.
One thing I'd gently flag as the failure mode, since you're clearly sold on the upside: the same mechanism that detects theory-fit can also detect mere similarity-to-us. The risk is selecting for people who already sound like the community rather than those who'd genuinely strengthen the theory — fit curdling into monoculture. Do you think AAI's setup has anything that guards against that, or is it a real exposure?
hiring is based on a months-long trial project and Jeremy reviews all the person's work so I think there's no real exposure beyond the fact that no one is perfect.
That's a strong answer, and it actually closes the loophole more tightly than I expected.
The months-long trial defeats the monoculture risk for a precise reason: you're no longer judging whether they sound like us, you're watching what they actually do to the theory over time. A short interview rewards similarity-signaling — saying the right shibboleths. A long project rewards genuine contribution, because you can't fake months of real work that strengthens the vision. Surface-fit and deep theory-fit, which a quick process conflates, get separated by time.
And Jeremy reviewing all the work is the other half — it keeps the sensing concentrated in a theory-holder rather than diffusing into "does the group like them." That sidesteps the committee-averaging failure where fit decays into consensus comfort.
So the real residual exposure is just the one you named — single-evaluator bandwidth and the fact that no judgment is perfect. Which is a very different, much smaller risk than structural monoculture.
On the Theory Building View the primary result of the programming activity is the theory held by the programmers. Since this theory by its very nature is part of the mental possession of each programmer, it follows that the notion of the programmer as an easily replaceable component in the program production activity has to be abandoned. Instead the programmer must be regarded as a responsible developer and manager of the activity in which the computer is a part. In order to fill this position he or she must be given a permanent position, of a status similar to that of other professionals, such as engineers and lawyers, whose active contributions as employers of enterprises rest on their intellectual proficiency.
were programmers consider less than engineers and lawyers in Naur's time?
Yes — and the gap was real and quite stark in 1985.
A few things were true then that make Naur's comparison pointed:
- Programming wasn't yet a "profession" in the formal sense engineering and law were. Those fields had licensure, professional bodies, codes of conduct, and protected status. Programming had none of that — no equivalent gatekeeping or recognized standing.
- It carried a clerical/operational legacy. Early "coding" was often seen as the lowly mechanical step below the real intellectual work of analysis and design — and historically much of it had been low-status, even feminized labor in the early computing era. The systems analyst was esteemed; the coder was a translator of someone else's design.
- The industrial-production framing Naur attacked in the previous paragraph was the managerial common sense: programmers as interchangeable line-workers, controllable by method and procedure, needing "only brief education."
So when Naur asks for status "similar to engineers and lawyers," he's not stating the obvious — he's making a provocative status claim against the prevailing managerial view. His ground for it is specific: not credentials or licensure, but that their contribution rests on intellectual proficiency — the theory — exactly like those professions.
There's a nice irony worth noting: the very thing that made it hard to grant programmers that status — the intangibility of what they know — is precisely what Naur has spent the whole essay arguing is the most valuable thing. What looked like "just typing code" was the lossy surface of an unwritten theory.
The raising of the status of programmers suggested by the Theory Building View will have to be supported by a corresponding reorientation of the programmer education. While skills such as the mastery of notations, data representations, and data processes, remain important, the primary emphasis would have to turn in the direction of furthering the understanding and talent for theory formation. To what extent this can be taught at all must remain an open question. The most hopeful approach would be to have the student work on concrete problems under guidance, in an active and constructive environment.
Conclusions
Accepting program modifications demanded by changing external circumstances to be an essential part of programming, it is argued that the primary aim of programming is to have the programmers build a theory of the way the matters at hand may be supported by the execution of a program. Such a view leads to a notion of program life that depends on the continued support of the program by programmers having its theory. Further, on this view the notion of a programming method, understood as a set of rules of procedure to be followed by the programmer, is based on invalid assumptions and so has to be rejected. As further consequences of the view, programmers have to be accorded the status of responsible, permanent developers and managers of the activity of which the computer is a part, and their education has to emphasize the exercise of theory building, side by side with the acquisition of knowledge of data processing and notations.
References
Brooks, R. E. Studying programmer behaviour experimentally. Comm. ACM 23(4): 207–213, 1980.
Feyerabend, P. Against Method. London, Verso Editions, 1978; ISBN: 86091–700–2.
Floyd, C. Eine Untersuchung von Software–Entwicklungs–Methoden. Pp. 248–274 in Programmierumgebungen und Compiler, ed H. Morgenbrod and W. Sammer, Tagung I/1984 des German Chapter of the ACM, Stuttgart, Teubner Verlag, 1984; ISBN: 3–519–02437–3.
Kuhn, T.S. The Structure of Scientific Revolutions, Second Edition. Chicago, University of Chicago Press, 1970; ISBN: 0–226–45803–2.
Medawar, P. Pluto’s Republic. Oxford, University Press, 1982: ISBN: 0–19–217726–5.
Moher, T., and Schneider, G. M. Methodology and experimental research in software engineering, Int. J. Man–Mach. Stud. 16: 65–87, 1. Jan. 1982.
Oskarsson, Ö Mechanisms of modifiability in large software systems Linköping Studies in Science and Technology, Dissertations, no. 77, Linköping, 1982; ISBN: 91–7372–527–7.
Polya, G. How To Solve It . New York, Doubleday Anchor Book, 1957.
Polya, G. Mathematics and Plausible Reasoning. New Jersey, Princeton University Press, 1954.
Popper, K. R., and Eccles, J. C. The Self and Its Brain. London, Routledge and Kegan Paul, 1977.
Ryle, G. The Concept of Mind. Harmondsworth, England, Penguin, 1963, first published 1949. Applying "Theory Building"
Applying “Theory Building”
Viewing programming as theory building helps us understand “metaphor building” activity in Extreme Programming (XP), and the respective roles of tacit knowledge and documentation in passing along design knowledge.
The Metaphor as a Theory
Kent Beck suggested that it is useful to a design team to simplify the general design of a program to match a single metaphor. Examples might be, “This program really looks like an assembly line, with things getting added to a chassis along the line,” or “This program really looks like a restaurant, with waiters and menus, cooks and cashiers.”
If the metaphor is good, the many associations the designers create around the metaphor turn out to be appropriate to their programming situation.
That is exactly Naur’s idea of passing along a theory of the design.
If “assembly line” is an appropriate metaphor, then later programmers, considering what they know about assembly lines, will make guesses about the structure of the software at hand and find that their guesses are “close.” That is an extraordinary power for just the two words, “assembly line.”
The value of a good metaphor increases with the number of designers. The closer each person’s guess is “close” to the other people’s guesses, the greater the resulting consistency in the final system design.
Imagine 10 programmers working as fast as they can, in parallel, each making design decisions and adding classes as she goes. Each will necessarily develop her own theory as she goes. As each adds code, the theory that binds their work becomes less and less coherent, more and more complicated. Not only maintenance gets harder, but their own work gets harder. The design easily becomes a “kludge.” If they have a common theory, on the other hand, they add code in ways that fit together.
An appropriate, shared metaphor lets a person guess accurately where someone else on the team just added code, and how to fit her new piece in with it.
Tacit Knowledge and Documentation
The documentation is almost certainly behind the current state of the program, but people are good at looking around. What should you put into the documentation?
That which helps the next programmer build an adequate theory of the program.
This is enormously important. The purpose of the documentation is to jog memories in the reader, set up relevant pathways of thought about experiences and metaphors.
This sort of documentation is more stable over the life of the program than just naming the pieces of the system currently in place.
The designers are allowed to use whatever forms of expression are necessary to set up those relevant pathways. They can even use multiple metaphors, if they don’t find one that is adequate for the entire program. They might say that one section implements a fractal compression algorithm, a second is like an accounting ledger, the user interface follows the model-observer design pattern, and so on.
Experienced designers often start their documentation with just
- The metaphors
- Text describing the purpose of each major component
- Drawings of the major interactions between the major components
These three items alone take the next team a long way to constructing a useful theory of the design.
The source code itself serves to communicate a theory to the next programmer. Simple, consistent naming conventions help the next person build a coherent theory. When people talk about “clean code,” a large part of what they are referring to is how easily the reader can build a coherent theory of the system.
Documentation cannot—and so need not—say everything. Its purpose is to help the next programmer build an accurate theory about the system.
HN Discussions
Amazing read! can you briefly summarize in short sentences in bullet points all the related topics we visited in our discussions? Include citations when relevant
Here's the same summary with references appended where our doc supports them:
- "The program lives in people's heads, not the code" is basically Daoism's "the Dao that can be spoken is not the real Dao"
- Docs are a lossy trace, which led us to Medawar's "the scientific paper is a fraud" and what's missing from LLM training data — Medawar, Pluto's Republic, 1982 (from Naur's reference list)
- Program-text decay maps onto the modern split of technical, cognitive, and intent debt and how AI shifts them — Fowler 2026, RDEL #137
- The jab at "scientific method" took us to Feyerabend on Galileo, the tower argument, and the untrustworthy telescope — Feyerabend, Against Method, 1978 (from Naur's reference list)
- Simplicity is Theory-relative, so LoC and cyclomatic complexity miss the point
- "Learning from the world" connects to Silver's new "Ineffable Intelligence" and its bet against human data — Wikipedia, CNBC, TechCrunch
- Program life/death/revival raises whether an AI with a world model and memory could join the crew
- Theory transfer needs close contact so hiring is theory-detection, à la the AAI community-and-trial approach
- Doing-under-guidance is why high-friction beats frictionless pipelines for building understanding
- "No right method" criticizes waterfall and Agile methods alike
- The Ship of Theseus as the right picture of a program's identity over time
Syntopical Read
Great read! can you compile below the list of books / papers we mentioned for Syntopical Reading?
Here's the list, grouped by the threads we followed:
The philosophy-of-science cluster (your syntopical set)
- Thomas Kuhn — The Structure of Scientific Revolutions (1962; 2nd ed. 1970) — paradigms, normal science, revolutions
- Karl Popper — conjecture & refutation, falsifiability (his The Self and Its Brain, with Eccles, 1977, is the one Naur actually cites for "World 3")
- Imre Lakatos — research programmes (hard core + protective belt)
- Paul Feyerabend — Against Method (1975/1978) — "anything goes," Galileo
- Peter Medawar — Pluto's Republic (1982) — "Is the scientific method a fraud?", imagination + criticism
That's the core set. The four-way dispute (Popper ↔ Kuhn ↔ Lakatos ↔ Feyerabend) is the natural spine for the syntopical exercise, with Ryle and Medawar as strong supporting voices.
which one should I buy from lakatos? and from popper?
web_search(type="search", queries="['Cambridge University Press Lakatos The…", query="Cambridge University Press Lakatos The M…")→"Server tool call executed."
{
"id": "ws_0fda338b74e02f83016a3faccdc1248191a0bbe7bb7775895d",
"name": "web_search",
"args": {
"type": "search",
"queries": "['Cambridge University Press Lakatos The Methodology of Scientific Research Programmes Philosophical Papers Volume 1', 'Routledge Popper Conjectures and Refutations The Growth of Scientific Knowledge', 'Routledge Popper The Logic of Scientific Discovery']",
"query": "Cambridge University Press Lakatos The Methodology of Scientific Research Programmes Philosophical Papers Volume 1"
},
"result": "Server tool call executed.",
"server": true
}
web_search(type="open_page", url="https://www.routledge.com/Conjectures-an…")→"Server tool call executed."
{
"id": "ws_0fda338b74e02f83016a3facdf38548191b08fc158454b31f8",
"name": "web_search",
"args": {
"type": "open_page",
"url": "https://www.routledge.com/Conjectures-and-Refutations-The-Growth-of-Scientific-Knowledge/Popper/p/book/9780415285940"
},
"result": "Server tool call executed.",
"server": true
}
Buy these:
-
Lakatos: The Methodology of Scientific Research Programmes: Philosophical Papers, Volume 1 This is the relevant one for the Popper/Kuhn/Feyerabend conversation. Cambridge lists chapters like “Falsification and the methodology of scientific research programmes,” “Popper on demarcation and induction,” and “Why did Copernicus’s research programme supersede Ptolemy’s?” — exactly your thread. Source: Cambridge University Press
-
Popper: Conjectures and Refutations: The Growth of Scientific Knowledge This is the best first Popper for your syntopical read. It is broader and more readable than The Logic of Scientific Discovery, and Routledge describes it as one of Popper’s clearest accessible statements of knowledge growing by trial and error. Source: Routledge
I would not start with Popper’s The Logic of Scientific Discovery unless you want the more technical/original version of falsifiability. And I would not buy Lakatos’s Proofs and Refutations for this specific project — it’s fascinating, but more about mathematical discovery than the philosophy-of-science debate around Kuhn/Popper/Feyerabend.
For inspectional reading, start with:
- Popper: “Science: Conjectures and Refutations”
- Lakatos: “Falsification and the Methodology of Scientific Research Programmes”
$0.0127
total=77,912 | in=76,150 | out=1,762 | cached=2.5% | reasoning=1,371 | $0.0127 | gpt-5.5
what is this about compared to the one you recommended? The Logic of Scientific Discovery
Compared with Popper’s Conjectures and Refutations, The Logic of Scientific Discovery is the more technical, foundational book.
- The Logic of Scientific Discovery is Popper’s original major work on the philosophy of science.
- Its central problem is: what separates science from non-science?
- Popper’s answer: scientific theories are not verified by evidence; they are scientific because they are falsifiable — they risk being proven wrong.
- It attacks the idea that science proceeds by simple induction from observations.
- It spends a lot of time on probability, testing, corroboration, and the logical structure of scientific theories.
By contrast:
- Conjectures and Refutations is more accessible.
- It gives the same big idea in a broader, essayistic way: knowledge grows through bold guesses and criticism.
- It is better for entering the Kuhn / Lakatos / Feyerabend conversation quickly.
- It is less formal and more readable.
So:
| Book | Best for |
|---|---|
| Conjectures and Refutations | First Popper; syntopical reading; grasping the debate |
| The Logic of Scientific Discovery | Deeper technical foundation of falsifiability |
For your current project, I’d still start with Conjectures and Refutations. Treat The Logic of Scientific Discovery as the “go deeper into Popper” book.
$0.0101
total=64,520 | in=63,850 | out=670 | cached=10.2% | reasoning=349 | $0.0101 | gpt-5.5
I found only volume 1 of 'Falsification and the Methodology of Scientific Research Programmes' in amazon
web_search(type="search", queries="['Lakatos The Methodology of Scientific …", query="Lakatos The Methodology of Scientific Re…")→"Server tool call executed."
{
"id": "ws_07dab82b2c315f5d016a3fae3970208191a7291a6236a25f90",
"name": "web_search",
"args": {
"type": "search",
"queries": "['Lakatos The Methodology of Scientific Research Programmes Volume 1 Cambridge University Press ISBN', 'Lakatos Falsification and the Methodology of Scientific Research Programmes Volume 1']",
"query": "Lakatos The Methodology of Scientific Research Programmes Volume 1 Cambridge University Press ISBN"
},
"result": "Server tool call executed.",
"server": true
}
web_search(type="open_page", url="https://www.cambridge.org/core/books/met…")→"Server tool call executed."
{
"id": "ws_07dab82b2c315f5d016a3fae4d70388191b7a03999669416fd",
"name": "web_search",
"args": {
"type": "open_page",
"url": "https://www.cambridge.org/core/books/methodology-of-scientific-research-programmes/8DBCEFE34A59BAD3D393FB958A4DC5FC"
},
"result": "Server tool call executed.",
"server": true
}
Yes — Volume 1 is exactly the one you want.
Small title clarification:
- The book is: The Methodology of Scientific Research Programmes: Philosophical Papers, Volume 1
- The main essay/chapter inside it is: “Falsification and the methodology of scientific research programmes”
Cambridge lists that chapter as Chapter 1, pp. 8–101, inside The Methodology of Scientific Research Programmes (Cambridge chapter page). The book page also says Volume I contains his papers on the philosophy of the physical sciences, while Volume II is more about philosophy of mathematics and related essays (Cambridge book page).
So for your syntopical reading, buy:
Imre Lakatos — The Methodology of Scientific Research Programmes: Philosophical Papers, Volume 1 Paperback ISBN: 9780521280310
No need to buy Volume 2 for this particular Popper/Kuhn/Feyerabend thread.
$0.0125
total=77,775 | in=76,449 | out=1,326 | cached=3.2% | reasoning=1,016 | $0.0125 | gpt-5.5
can you gather a bit the thoughts, points & discussion to guide my syntopical reading exploration?
Here's a synthesis to orient your syntopical reading. I'll keep it to the structure first, and we can deepen any part you like.
The shared question (your neutral terms)
All these authors are circling one question: what makes knowledge-building rational, if not a fixed method? Frame it neutrally so no single author's vocabulary dominates — e.g. "the role of explicit rules vs. tacit judgment in producing reliable knowledge."
The four positions on the spine
- Popper — rules can be stated: conjecture boldly, then try to falsify. Rationality = criticism against a logical standard.
- Kuhn — in practice scientists don't falsify; they puzzle-solve inside a paradigm learned through exemplars, until crisis flips it. Norms are paradigm-relative.
- Lakatos — rescues a softened Popper: judge research programmes (hard core + protective belt) as progressive or degenerating over time, not single experiments.
- Feyerabend — no rule survives the history of real breakthroughs; "anything goes." The provocateur.
The connecting thread to Naur
The reason all this sat under a programming essay: Naur's "theory" is the tacit, similarity-based, non-rule-governed understanding — and he recruits these philosophers to argue that even science runs on that, not on method. So your reading isn't a detour from Naur; it's the philosophical foundation of his central claim.
Want me to lay out the recurring sub-themes next (the questions you'd actually track across all four as you read) — that's the real engine of a syntopical reading?
if the sub-themes were discussed yes, if not no
Yes — several genuinely surfaced in our discussion, so here they are (only the ones we actually touched):
- Tacit judgment vs. statable rules — Naur's faces/tunes/wine; Feyerabend's "no rule survives." Track: does each author think the essential skill can be written down?
- Theory-laden observation — Galileo's tower argument and telescope; we said data isn't clean bedrock. Track: how much does each let background theory shape what counts as evidence?
- Exemplars vs. abstract laws — Kuhn's 187ff (paradigms as solved examples), which Naur borrowed. Track: is competence built from examples or from principles?
- The imaginative spark vs. the critical loop — Medawar's two-stroke engine; our LLM/Sutton tangent. Track: where does each locate discovery — in the guess, or in the testing?
- Does science have special standing? — the Feyerabend overreach most reject. Track: rationality as universal, paradigm-relative, or none.