How often do variables in dynamically-typed code actually change type?
I've decided to self-archive some of my Stack Exchange answers that contain real research work, particularly the surveys that represent a lot of synthesis of the literature, so that I've got them here even if Stack Exchange goes away and I can extend and annotate them further. This investigation grew from an answer I wrote to a Stack Exchange question Are there metrics on how often variables in dynamically-typed languages change their type (not "parametrically") about programmers in dynamically-typed languages using the same variable with multiple data types, outside of generic functions and automated type coercions.
A lot of work on type inference for dynamically-typed languages faces this issue, and often enough in practice that it comes up as something to be dealt with, while rarely enough that it's sometimes explicitly excluded from support. Some work has looked specifically at the use of dynamically-typed features on their own.
In most cases the relevant results are on the side of something else, either an inference engine or analysis of more complex dynamic features. It doesn't seem like anyone has looked specifically at reassignment, which I do find a little surprising.
Further down I'll note some limitations, and other cases that may or may not be in scope here. Exactly what it means for a variable to hold a new type is potentially a bit tricky to nail down. For example, what constitutes a type in these languages? is it necessary that the variable be reassigned?
Xia et al. performed an empirical study on dynamic-typing behaviours in Python programs on a substantial corpus of popular real Python systems, finding that at least 6.9% of identifiers had multiple types, and another 13.4% were undetermined by their methodology, while at least 79.7% of identifiers had only a single type. They show that even assigning two different types of literal to a variable sequentially has real incidence, while still being uncommon. They present a heat map of this matrix in Figure 5, but unfortunately an "expression" category squashes together what could be many different kinds of conversion into one (and I think the table over-eagerly aligns changes of type with changes of expression kind).
Chen et al. performed another Python study and found an average of 311 instances of variables being given different types across their benchmark. I don't think this average metric they reported is particularly useful, but all but one of the systems under study had more than 50, and these were the most common dynamic-typing behaviours they measured by a large margin. Excluding the low and high outliers, between 10% and 25% of methods they examined in each system contained one or more of the traits they examined, and these would primarily be variable typing. However, most instances found did combine string with another type, noting that
It is often the case that the value of a variable is parsed from the user inputs and its type is determined by the source of the inputs (e.g., XML files, database or command line).
They do present examples from real-world systems that show other replacements of values with different types, however, including strings. An example is given from IPython where a variable holding a dictionary is assigned one of the values from that dictionary. The reporting does not have enough detail to distinguish all of the cases desired to rule out.
Furr et al. studied Ruby programs using a profile-guided tool to infer static types for unannotated code. To produce typecheckable programs, the tool required refactoring of the benchmark programs to remove precisely this issue: they performed 11 refactorings to break multi-typed variables into multiple uni-typed variables, out of 226 total refactorings across their suite (a number of these related to specific limitations of this version of the tool, like support for dynamic type tests, so the "true" proportion would be higher). They also note 12 instances of fundamentally untypeable code, notably within the optparse module that parses command-line options: while these issues are string->other conversions, the problem is that the repeated control flow gives variables multiple values for different options, so you might count these as well. However, both classes are relatively rare and do not occur in the majority of their benchmark programs.
Pradel et al. produced a tool for finding type inconsistencies in JavaScript code, and benchmarked it against a suite of published code. This system was focused on finding potential bugs, rather than observing idioms, but it did encounter several cases of mixed types within functions that seem intentional; most of these involve the "undefined" type and another, which is meaningful within JavaScript's type model but could be analysed as a single nullable type as well. Other combinations also existed, but were less common.
I believe one of the Vitek analyses of R also touches on this question, but I haven't been able to find which one, if it does exist. I have seen a presentation on the incidence of this pattern in PHP code (quite common), but can't find any archival publication of the result; it's an explicit motivator for much of the Hack work, so I expect it has been measured to matter on the Facebook codebase too.
One limitation that all of these face is that they are benchmarking against published code, which may tend not to use some of these features as much. It is likely that this overwriting is much more common in interactive use, which is inherently ephemeral and so not included in benchmarks. For example, a significant amount of R usage is entirely interactive, and maintaining a single variable for the in-progress results of where the user is up to is not uncommon, notwithstanding that the variable may hold different types internally as further analysis steps are run (and these different types may or may not be significant or known to the user). Some of this sort of interactive use may persist into non-published scripts calcified from interactive sessions.
It's also the case, though, that code where a variable has type X up to a point, and type Y thereafter, isn't necessarily resting on dynamic typing at all: without loss of generality, this can be taken as two separate variables that have the same name, one shadowing the other. It's only cases where the variable is accessed from a loop or through lexical capture, or the type changes conditionally, where the ability to mix types together really matters. This is what the refactoring from the PRuby work above rested on. The reverse can also be true: idiomatic JavaScript will declare uninitialised variables before a loop to be set inside, such as in a search — but this is strictly a change of type within JavaScript's type model, because "undefined" is its own type! Similarly, Python None is used in the same role, and is not a bottom type either. It's necessary to drill down very specifically into what is meant by type changes in order to quantify them.
I haven't touched here on another sort of "type" change: meta-mutable object values may be seen to have a different type each time one of their properties is added or removed. This could happen either as a value is built up originally (consider idiomatic JavaScript let x = {}; x.a = 1; x.b = 2;), either inline or by passing it to other code to populate; it can also happen long afterwards while the variable is still in scope somewhere. A change could also arise from modifications made to inheritance parents. Richards et al. showed that all of these sorts of change are very common in JavaScript code, and many are also possible and idiomatic in other languages on the more dynamic end of the spectrum. In Ruby, both monkeypatching existing types and eigenclass modifications are normal parts of using the language. These would be changes of a variable's type from some perspectives, and not others, and I'm not sure whether they're in scope of this question or not.
References
- , , , , and . . “An Empirical Study on Dynamic Typing Related Practices in Python Systems”. In Proceedings of the 28th International Conference on Program Comprehension (ICPC '20): 83–93. ACM, New York, NY, USA. https://doi.org/10.1145/3387904.3389253.
- , and . . “Profile-guided static typing for dynamic scripting languages”. In Proceedings of the 24th ACM SIGPLAN conference on Object oriented programming systems languages and applications (OOPSLA09): 283–300. ACM, New York, NY, USA. https://doi.org/10.1145/1640089.1640110.
- , and . . “TypeDevil: Dynamic Type Inconsistency Analysis for JavaScript”. In 2015 IEEE/ACM 37th IEEE International Conference on Software Engineering (ICSE): 314–324. IEEE. https://doi.org/10.1109/icse.2015.51.
- , , and . . “An analysis of the dynamic behavior of JavaScript programs”. In Proceedings of the 31st ACM SIGPLAN Conference on Programming Language Design and Implementation (PLDI '10): 1–12. ACM, New York, NY, USA. https://doi.org/10.1145/1806596.1806598.
- , , , and . . “An Empirical Study of Dynamic Types for Python Projects”. In Lecture Notes in Computer Science (SATE 2018): 85–100. Springer International Publishing, Cham. ISBN: 9783030042714. https://doi.org/10.1007/978-3-030-04272-1_6.