Theory in the Natural Sciences: A View from Chemistry

Author: Dclean
Reviewers: 黍离, 白書, and 一毫秒的永恒

  Human understanding of nature begins with observing natural phenomena. The immediate impressions left by those observations became some of the earliest “knowledge” in human history. In pursuit of deeper understanding, people gradually classified and organized that knowledge, connecting it through logic. From this process came the first rudiments of “theory.”

  This article considers how science understands theory, the role of theory in the natural sciences, and the relationship between theory and practice.

Part I | How Disciplines Differ: Subject Matter and Method

  We often ask what truly distinguishes physics from chemistry. We also hear claims that chemistry is merely a branch of physics. To think clearly about such questions, we must first identify what separates one discipline from another.

  The most obvious distinction is subject matter. The humanities, for example, study patterns in the development of human society, a fundamentally different undertaking from the natural sciences’ study of how the natural world works. Yet subject matter is not the whole story. Once a discipline has an object of study, it also needs methods suited to that object. Subject matter and method, in fact, select each other. The separation of physics and chemistry offers a useful example of this reciprocal selection.

  Physics and chemistry are both natural sciences concerned with nonliving systems, and at first glance their subject matter may not seem very different. In secondary school, you may have read a statement such as: “Chemistry studies the properties of matter from the atomic level up to the supramolecular scale.” But that definition raises an obvious question. Why should physics cover both the very large—from macroscopic to cosmic scales—and the very small, below the atom, while chemistry is confined to the territory in between?

  Let us introduce the “unit of observation”: the most basic part of the objective reality involved in a problem that can be treated as a whole. What are the usual units of observation in elementary physics and elementary chemistry?

  At the microscopic scale, elementary physics studies interactions among particles. Individual particles are plainly the units of observation: we can write equations for a particle or a system of particles and solve for its behavior. The same is true at cosmic scales. We treat individual celestial bodies as wholes, study their interactions, and write equations for a body or a system of bodies to determine its behavior.

  In a chemical reaction, however, no comparable unit of observation presents itself. A molecule cannot serve because it does not remain unchanged during the reaction. Nor can an atom, because atoms gain and lose electrons. A chemical reaction offers no unit that both remains unchanged and has an internal structure we can ignore for the problem at hand. The mathematical equations describing a chemical system therefore cannot be reduced to the simplicity found in elementary physics. Early in the history of science, this was the highest degree of complexity that the mathematics then available could handle. For the same reason, secondary-school chemistry cannot approach its problems with the rigorous mathematical methods familiar from physics. It must remain at the macroscopic level and rely on comparatively empirical theories.

  In this sense, the fundamental difference between physics and chemistry arises from the nature of their objects: subject matter determines method. From the perspective of scientific history, however, the direction can also be reversed. By observing nature, people gradually developed scientific methods of inquiry. Once a method suited to a certain kind of object had emerged, inquiry shifted from unstructured observation to observation guided by that method. Human observation is necessarily limited, yet methods derived from a finite set of observed objects must later be applied to objects never seen before. In that process, each distinct method finds the subject matter to which it is best suited. Selection then runs in the opposite direction. This two-way selection between subject matter and research method marks the beginnings of disciplinary differentiation.

Part II | The Birth of Theory and the Development of a Discipline

  The first theories in the natural sciences arose when observational results were made logical and systematic. Theory lifted knowledge from simple, intuitive recognition into an organized logical structure. Its defining property is logical coherence, and that is what distinguishes a theory from an empirical conclusion.

  The earliest theories were induced from experimental findings and tested against further experiments. The soundness of this approach is readily apparent, and it has persisted across every field and period of research. Consider the training and test sets used in today’s popular field of machine learning. To find a relationship between input variables and a target variable, an algorithm fits the training data. Researchers then use the test set to evaluate whether that fitted relationship holds over a broader range. If test inputs passed through the fitted relationship produce target values close to reality, the learning has been relatively successful. The relationship produced by machine learning is not itself a “theory,” however; it is an “empirical conclusion,” even when its mathematical or logical form is clear and rigorous. A computer algorithm fitted the relationship. It has no inherent physical meaning, nor can it necessarily serve as a starting point from which to deduce new conclusions logically. Theories, by contrast, can provide major and minor premises for each other, logically generate new theories, and be derived from or corroborated by others within the same system. This is what their “logical character” means. If researchers proceed from a mathematical relationship to a clear physical picture, derive a formula with definite physical meaning, and integrate it into an existing theoretical system, the result has risen to the level of theory.

  Later theories need not arise directly from experimental data; they can also be deduced logically from theories already in place. This route is especially important in the abstract sciences, most notably in mathematics and its “axiomatic systems.” A mathematical axiomatic system still differs fundamentally from a theoretical system in natural science because the two study different kinds of objects. This is why I emphasized the logical character of theory above.

  Logic also plays a vital role in the growth of theory and the maturation of a discipline. The history of natural science is the history of scientific theories being born, developing, and flourishing. Natural science is defined not only by its study of nature, but also by the systematic organization of its theories and the logical connections among them. Only a systematic body of theory can support a sound mechanism for detecting and correcting errors, and such a mechanism is essential to disciplinary progress. Its opposite is the mere accumulation of facts. Knowledge gathered as disconnected scraps remains fragmented; it is difficult to learn and even harder to test for correctness. This is the weakness of an empirical discipline. Traditional Chinese medicine, long a subject of debate, offers an example. Because it was established and developed so long ago, researchers had little means of exchanging information and could do little beyond recording what they saw, heard, and obtained from experiments. Over centuries, those records accumulated into a system. Yet every experiment is a chaotic system. At every link in such an empirical chain, individual variation or differences in conditions may make a result unreliable. Errors accumulated over a thousand years do not correct themselves; they compound. The result is what we see today: a medical book may contain hundreds or thousands of ineffective prescriptions alongside many effective ones, because both correct and incorrect experimental conclusions accumulated over time. This flaw is fatal to a discipline’s future. A sound theoretical system cannot emerge while the body of knowledge contains too many errors in practice. That is why modern science places such importance on reproducibility. If every experimental claim were accepted without independent replication, random errors would accumulate until the discipline’s theoretical structure collapsed.

  The systematic, logical character of theory also matters to learners. It determines how they grasp the core of a body of knowledge. Simply stating a theory has a very different effect from explaining how it was established and guiding learners to derive it for themselves. In my view, the best way to learn a field is to follow the path along which its theoretical system was built. A few disciplines developed by unusual routes that do not match ordinary patterns of thought, but such cases are rare. This is the value of studying the history of science. By retracing the construction of a theoretical system, the learner partly takes on the researcher’s perspective. The researcher sees far more: how a theory arose and how it generated new theories. Following that process can give the learner a much deeper understanding. Only a systematic logical structure makes this possible. A coherent body of knowledge resembles a branching tree. Start with the trunk and grasp the core, and you can understand the field deeply enough to learn its branches more effectively. A heap of unrelated facts resembles a pile of bricks. Stacking them one by one reveals neither their deeper content nor their logical relationships.

  As an aside, the natural sciences may now have highly developed theoretical systems, but learners encounter those systems through educational resources. If a resource lacks logic and turns theory back into a pile of facts, the advantage of the theoretical system is lost.

  How accessible a discipline is to learners directly affects the pace of its development. This is one reason the natural sciences that flourish today all have well-developed theoretical systems, while empirical sciences have gradually moved to the margins.

Part III | How Theory and Experiment Divide the Work

  A saying has circulated widely online in recent years: “Theory is when you know everything but nothing works. Practice is when everything works but no one knows why. In our lab, theory and practice are combined: nothing works and no one knows why.” A version closer to its intended meaning would be: “Theoretical work understands every principle but has no practical value. Practical work produces results that work, but no one can explain why. Our laboratory combines theory and practice: none of our results is useful, and no one knows why.” The last line is pure self-mockery and can be set aside. The first two are exaggerated, but they contain some truth.

  My supervisor once told me that the ideal of computational chemistry is to guide experiments. How? With a sound theoretical framework and accurate calculations, theory can point out the right path before an experiment begins. Yet that ideal cannot always be achieved. Our research is often wholly disconnected from practice, but that does not make it meaningless. Accumulating such results adds brick after brick to theoretical science. A paper published several years ago in Science by my institute examined only the dynamics of a chemical reaction involving three atoms. The work was completely removed from practical application, yet it was excellent research worthy of Science. Theoretical work in every new field begins with the simplest cases and gradually develops until it can address real systems.

  Theory works because its principles are reasonable. I use “reasonable” rather than “correct” because no existing theory can be declared conclusively true. A theory is only reasonable at humanity’s current level of understanding; a sound approximation to a theory also counts as reasonable here. Only a reasonable theory can produce conclusions close to the facts, and that is the basis for testing theory against reality. This demand for reasonableness is both the strength and the weakness of theoretical research. Its strength is that, given a sound theory, one can eliminate every other interfering factor and obtain a completely “clean,” fully reproducible result—something no experiment can achieve. Its weakness is the point just made: a sound result requires a sound theory in advance.

  Experiment is almost the reverse. Its weakness is that, no matter how carefully conditions are controlled, even the best control-variable design cannot make every irrelevant variable identical. An experiment is therefore a chaotic system, and its results are less reproducible than theoretical ones. Researchers often publish a paper only to receive an email saying that its method cannot be replicated. Investigation may uncover a trace impurity on an instrument, or even a minute difference in impurities because the reagents came from another manufacturer. The strength of experiment is that any completed experiment produces a result. The reason need not be known beforehand, and the result is definite under the conditions of the experiment. Experimenters do not—and cannot—know those conditions completely or reproduce them exactly. They know only the few macroscopic variables they can control. In the examples above, they controlled many variables but did not anticipate a tiny residue left on an instrument. Even without any understanding of the underlying principles, experiments can produce exciting and important discoveries. Theory cannot do that. Empirical conclusions differ from theories in this respect as well: many predict successfully even when their mechanisms remain unknown.

  The complementary strengths of theory and experiment naturally allow them to support each other in actual research. If an experiment disagrees with a theoretical prediction and problems in the theory have been ruled out, the experimental conditions should be questioned. This can be valuable when tracing experimental errors. Conversely, if a theory’s predictions consistently diverge from reality across a particular domain, the theory probably does not apply there and should be improved or replaced.

  In general, a healthy discipline develops theory ahead of practice. Experimental science advances rapidly under theoretical guidance. In a discipline where applications outpace theory, the frontier often reaches a bottleneck and advances only slowly.

The author, Dclean, studied in the Department of Chemistry at Nanjing University and conducted research at the Institute of Theoretical and Computational Chemistry.