What factors make a programming language more (or less) beginner-friendly?


  1. Victoria University of Wellington

Building from an answer I wrote to a Stack Exchange question What makes a language 'beginner-friendly'?, asking what aspects of a programming language made it more (or implicitly less) suitable for a beginner, here I discuss what's in the literature about this topic. My PhD and a lot of my subsequent work has been with Grace, intended as an educational language, but this is a more complex question than it seems on the surface.


One challenge with addressing this question is that the evidence base is much thinner than you might expect, and much of the wisdom floating around isn't based on anything much, and quite often does not hold up when examined — but so little is examined that we don't necessarily know which.

There just aren't as many actual experiments looking at programming-language design in total as you might expect — Antti-Juhani Kaijanaho's 2012 PhD thesis identifies somewhere between 37 and 137, depending on your inclusion criteria, and there haven't been vast numbers since then. Many of those are not specifically relevant to beginners.


The most significant recent study on language syntax and keywords for novices comes from Andreas Stefik and Susanna Siebert in 2013, An Empirical Investigation into Programming Language Syntax. Some of the results there are likely to be surprising, while others are obvious at least in retrospect. Overall, the major takeaways are that

  • Use of words rather than symbols generally helps matters, but ...
  • ... the typical words used in programming languages are not very effective.
    For example, terms like "foreach", "while", and "for" perform very poorly, while data type names like "float" and "string" don't do well either.
  • Metaphorical terms are especially weak: things like throw and catch do worse than "error" — but in place of both of those, which would probably be confusing too.
  • In general, using fewer tokens to express something performs better than using more, but this is in tension with using helpful non-symbolic terms.
  • Neither beginners nor more-experienced programmers are very consistent in just about any facet, except that the experienced programmers favour the choices they're accustomed to (for example, the term cout tested very favourably for output).

When experimenting on whole languages, Ruby, Python, and their Quorum language performed best. However, these experiments were still only in an artificial environment and did not measure learning, only accuracy.


There is a reasonable amount of evidence that static typing produces better results for novice programmers by catching their errors, but also that type annotations produce more syntax errors for them. The most well-known work in this vein is by Stefan Hanenberg, but it has been replicated a few times. Pedagogical experience reports often highlight the latter point as more significant at the very beginning — it's that "more tokens" issue again.

Block-based programming environments, which prevent the creation of various kinds of syntax and type error, have been shown to perform better with novices than textual languages even when the language is otherwise identical, such as by Thomas W. Price and Tiffany Barnes in 2015, Comparing Textual and Block Interfaces in a Novice Programming Environment; there has also been significant work by David Weintrop and others. However, while these provide good support in the early stages of learning, they can also impose meaningful exit friction when the novice advances beyond what the block environment gives them. Most block environments also provide very easy access to graphical or interactive primitives, and may show live stepping or other debugging affordances.

"Good" error messages do increase learner performance, but those that are "too helpful" are counterproductive, especially if they can make a bad guess about the user's intent. A substantial number of novices will "freeze" when there are many error or even warning indicators given to them, but presenting all errors at once outperforms one-at-a-time reporting. Allowing erroneous programs to run partially has shown some effectiveness. Improvements to error reporting have been studied among others by Becker, Brett A., Graham Glanville, Ricardo Iwashima, Claire McDonnell, Kyle Goslin, and Catherine Mooney in 2016, Effective Compiler Error Message Enhancement for Novice Programming Students. Negative responses to error messages perceived as directed at the user are common in new learners and the phrasing and content needs to be balanced carefully.

Limited sub-languages that progress towards the unrestricted language have shown value as far back as SP/k, but it is Racket more recently that has brought them to the fore. The advantage these systems provide is that learners do not "stumble" into advanced functionality they didn't intend to reach; this sometimes happens when error messages lead them astray (a common anecdote is a novice who turns their entire Java program static after they try to fix an error saying they are unable to access an instance method or field from main). However, these limited languages also introduce more kinds of error and more ways for the learner to go wrong.

Although the language "paradigm" seems like it should form a major component of the answer, the evidence for and against functional, imperative, object-oriented, procedural, ... models for learners is very mixed over time. There does not seem to be a strong recommendation that could be made in isolation backed by anything but personal assertion (and many have done so). The other aspects already mentioned, and any structured teaching wrapped around it, are much bigger factors.

Localisation is a significant issue, particularly for younger learners. Learning keywords and following already-opaque error messages in an unknown language is challenging and new programmers understandably tend to perform better when the language keywords, library elements, word order, and diagnostic messages match their native language. For beginners who are non-English speakers (or are also beginners there), a localised language of some kind will probably help them at first. However, in many cases those localised languages that do exist have limited progression paths out of them and limited library support. Hedy is a language that tries to navigate this, with detailed localisation including the complex elements like word and writing order, but representing Python behind the scenes, and there's been quite a bit of study of how that works out.

Finally, and this is a bit out of left field, the programming language that the most beginners pick up unassisted and to productive use is the spreadsheet, a spatial dataflow language. These do have very high error rates, but immediate accessibility that virtually no other programming environment matches. The true answer for a randomly-selected "beginner" is probably Excel.


Other topics I won't touch on more, but are worth keeping in mind also:

  • There is a concept in educational psychology called transfer. A beginner won't be a beginner forever, and if they soon have to move on to another language they likely will not be able to transfer what they've learned already easily, unless they have directed instruction guiding them. It's often assumed that if the new target is "similar enough" to what they have been using already then this will happen for free, but research does not bear that out. A language that is merely similar to one the learner may want to use in future may not be helpful.
  • In considering the starting question, "Which coding languages should a beginner learn?", there are a lot of other elements outside of the language designs themselves, and these are likely to dominate in practice. Those with significant communities or commercial appeal are often going to be better choices than those more ideal for a beginner to learn.
  • If the beginner is to pick up the language on their own, the answer is likely different to if it is to be taught to them. In particular, teaching with explicit bridging content can allow the "exit path" from novice-specific languages that is much harder to navigate alone.
  • Accessibility varies significantly by the tooling, more so than the language, but these are often very correlated. IDE support for screenreaders and other assistive technology varies greatly, and some language families have a much worse time than others.

Broad overviews of the long-standing research on learning programming are in

References

  • Becker, Brett A., Graham Glanville, Ricardo Iwashima, Claire McDonnell, Kyle Goslin and Catherine Mooney. . “Effective compiler error message enhancement for novice programming students”. In Computer Science Education 26 (2-3): 148–175. Informa UK Limited. https://doi.org/10.1080/08993408.2016.1225464.
  • Black, Andrew P., Kim B. Bruce, Michael Homer and James Noble. . “Grace: the absence of (inessential) difficulty”. In Proceedings of the ACM international symposium on New ideas, new paradigms, and reflections on programming and software (SPLASH '12): 85–98. ACM, New York, NY, USA. https://doi.org/10.1145/2384592.2384601.
  • European Spreadsheet Risk Interest Group. n.d. “Research and Best Practice”. Accessed 2024-04-03. Online.
  • Hermans, Felienne. . “Hedy: A Gradual Language for Programming Education”. In Proceedings of the 2020 ACM Conference on International Computing Education Research (ICER '20): 259–270. ACM, New York, NY, USA. https://doi.org/10.1145/3372782.3406262.
  • Homer, Michael. . “Graceful Language Extensions and Interfaces”. Victoria University of Wellington Library. https://doi.org/10.26686/wgtn.17008246.
  • Kaijanaho, Antti-Juhani. . “Evidence-based Programming Language Design: A Philosophical and Methodological Exploration (PhD thesis)”. Thesis. University of Jyväskylä. ISBN: 9789513963880. Online.
  • Pears, Arnold, Stephen Seidman, Lauri Malmi, Linda Mannila, Elizabeth Adams, Jens Bennedsen, Marie Devlin and James Paterson. . “A survey of literature on the teaching of introductory programming”. In ACM SIGCSE Bulletin 39 (4): 204–223. Association for Computing Machinery (ACM). https://doi.org/10.1145/1345375.1345441.
  • Price, Thomas W. and Tiffany Barnes. . “Comparing Textual and Block Interfaces in a Novice Programming Environment”. In Proceedings of the eleventh annual International Conference on International Computing Education Research (ICER '15): 91–99. ACM, New York, NY, USA. https://doi.org/10.1145/2787622.2787712.
  • Robins, Anthony, Janet Rountree and Nathan Rountree. . “Learning and Teaching Programming: A Review and Discussion”. In Computer Science Education 13 (2): 137–172. Informa UK Limited. https://doi.org/10.1076/csed.13.2.137.14200.
  • Stefik, Andreas and Susanna Siebert. . “An Empirical Investigation into Programming Language Syntax”. In ACM Transactions on Computing Education 13 (4): 1–40. Association for Computing Machinery (ACM). https://doi.org/10.1145/2534973.
  • Swidan, Alaaeddin and Felienne Hermans. . “A Framework for the Localization of Programming Languages”. In Proceedings of the 2023 ACM SIGPLAN International Symposium on SPLASH-E (SPLASH-E '23): 13–25. ACM, New York, NY, USA. https://doi.org/10.1145/3622780.3623645.
Kaijanaho, Antti-Juhani. . “Evidence-based Programming Language Design: A Philosophical and Methodological Exploration (PhD thesis)”. Thesis. University of Jyväskylä. ISBN: 9789513963880.
Stefik, Andreas and Susanna Siebert. . “An Empirical Investigation into Programming Language Syntax”. In ACM Transactions on Computing Education 13 (4): 1–40. Association for Computing Machinery (ACM).
Price, Thomas W. and Tiffany Barnes. . “Comparing Textual and Block Interfaces in a Novice Programming Environment”. In Proceedings of the eleventh annual International Conference on International Computing Education Research (ICER '15): 91–99. ACM, New York, NY, USA.
Becker, Brett A., Graham Glanville, Ricardo Iwashima, Claire McDonnell, Kyle Goslin and Catherine Mooney. . “Effective compiler error message enhancement for novice programming students”. In Computer Science Education 26 (2-3): 148–175. Informa UK Limited.
European Spreadsheet Risk Interest Group. n.d. “Research and Best Practice”.
Black, Andrew P., Kim B. Bruce, Michael Homer and James Noble. . “Grace: the absence of (inessential) difficulty”. In Proceedings of the ACM international symposium on New ideas, new paradigms, and reflections on programming and software (SPLASH '12): 85–98. ACM, New York, NY, USA.
Swidan, Alaaeddin and Felienne Hermans. . “A Framework for the Localization of Programming Languages”. In Proceedings of the 2023 ACM SIGPLAN International Symposium on SPLASH-E (SPLASH-E '23): 13–25. ACM, New York, NY, USA.
Hermans, Felienne. . “Hedy: A Gradual Language for Programming Education”. In Proceedings of the 2020 ACM Conference on International Computing Education Research (ICER '20): 259–270. ACM, New York, NY, USA.
Pears, Arnold, Stephen Seidman, Lauri Malmi, Linda Mannila, Elizabeth Adams, Jens Bennedsen, Marie Devlin and James Paterson. . “A survey of literature on the teaching of introductory programming”. In ACM SIGCSE Bulletin 39 (4): 204–223. Association for Computing Machinery (ACM).
Robins, Anthony, Janet Rountree and Nathan Rountree. . “Learning and Teaching Programming: A Review and Discussion”. In Computer Science Education 13 (2): 137–172. Informa UK Limited.
Homer, Michael. . “Graceful Language Extensions and Interfaces”. Victoria University of Wellington Library.