tup = (4, 5, 6)McKinney Chapter 3 - Built-In Data Structures, Functions, and Files
FINA 6333 for Spring 2025
Introduction
We must understand Python’s core functionality to use NumPy and pandas. Chapter 3 of McKinney (2022) discusses Python’s core functionality. We will focus on the following:
- Data structures
- tuples
- lists
- dicts (also known as dictionaries)
- List comprehensions
- Functions
- Returning multiple values
- Using anonymous functions
Note: Indented block quotes are from McKinney (2022) unless otherwise indicated. The section numbers here differ from McKinney (2022) because we will only discuss some topics.
Data Structures and Sequences
Python’s data structures are simple but powerful. Mastering their use is a critical part of becoming a proficient Python programmer.
Tuple
A tuple is a fixed-length, immutable sequence of Python objects.
We cannot change a tuple after we create it because tuples are immutable. A tuple is ordered, so we can subset or slice it with a numerical index. We will surround tuples with parentheses but they are not required.
Python is zero-indexed, so zero accesses the first element in tup!
tup[0]4
tup[1]5
tup[2]6
nested_tup = ((4, 5, 6), (7, 8))nested_tup[0](4, 5, 6)
nested_tup[0][0]4
tup = ('foo', [1, 2], True)If an object inside a tuple is mutable, such as a list, you can modify it in-place.
# tup[2] = False # gives an error, because tuples are immutable (unchangeable)tup[1].append(3)
tup('foo', [1, 2, 3], True)
You can concatenate tuples using the + operator to produce longer tuples:
Tuples are immutable, but we can combine two tuples into a new tuple.
(1, 2) + (1, 2)(1, 2, 1, 2)
(4, None, 'foo') + (6, 0) + ('bar',)(4, None, 'foo', 6, 0, 'bar')
Multiplying a tuple by an integer, as with lists, has the effect of concatenating together that many copies of the tuple:
This multiplication behavior is the logical extension of the addition behavior above. The output of tup + tup should be the same as that of 2 * tup.
('foo', 'bar') + ('foo', 'bar')('foo', 'bar', 'foo', 'bar')
('foo', 'bar') * 2('foo', 'bar', 'foo', 'bar')
Unpacking tuples
If you try to assign to a tuple-like expression of variables, Python will attempt to unpack the value on the righthand side of the equals sign.
tup = (4, 5, 6)
a, b, c = tupa4
b5
c6
(d, e, f) = (7, 8, 9) # the parentheses are optional but helpful!d7
e8
f9
We can unpack nested tuples!
tup = 4, 5, (6, 7)
a, b, (c, d) = tupTuple methods
Since the size and contents of a tuple cannot be modified, it is very light on instance methods. A particularly useful one (also available on lists) is count, which counts the number of occurrences of a value.
a = (1, 2, 2, 2, 3, 4, 2)
a.count(2)4
Python is zero-indexed!
a.index(2)1
List
In contrast with tuples, lists are variable-length and their contents can be modified in-place. You can define them using square brackets [ ] or using the list type function.
a_list = [2, 3, 7, None]
tup = ('foo', 'bar', 'baz')
b_list = list(tup)a_list[2, 3, 7, None]
b_list['foo', 'bar', 'baz']
Python is zero-indexed!
a_list[0]2
Concatenating and combining lists
Similar to tuples, adding two lists together with + concatenates them.
[4, None, 'foo'] + [7, 8, (2, 3)][4, None, 'foo', 7, 8, (2, 3)]
The .append() method adds its argument as the last element in a list.
xx = [4, None, 'foo']
xx.append([7, 8, (2, 3)])
xx[4, None, 'foo', [7, 8, (2, 3)]]
If you have a list already defined, you can append multiple elements to it using the extend method.
x = [4, None, 'foo']
x.extend([7, 8, (2, 3)])
x[4, None, 'foo', 7, 8, (2, 3)]
Check your output! It will take you time to understand all these methods!
Slicing
Slicing is very important!
You can select sections of most sequence types by using slice notation, which in its basic form consists of start:stop passed to the indexing operator [ ].
Recall that Python is zero-indexed, so the first element has an index of 0. A consequence of zero-indexing is that start:stop is inclusive on the left edge (start) and exclusive on the right edge (stop).
seq = [7, 2, 3, 7, 5, 6, 0, 1]
seq[7, 2, 3, 7, 5, 6, 0, 1]
seq[5]6
Python is zero-indexed, so left edge of slide is included and right edge is excluded!
seq[1:5][2, 3, 7, 5]
Either the start or stop can be omitted, in which case they default to the start of the sequence and the end of the sequence, respectively.
seq[:5][7, 2, 3, 7, 5]
seq[3:][7, 5, 6, 0, 1]
Negative indices slice the sequence relative to the end.
seq[-1]1
seq[-1:][1]
seq[-4:][5, 6, 0, 1]
seq[-4:-1][5, 6, 0]
seq[-6:-2][3, 7, 5, 6]
A step can also be used after a second colon to, say, take every other element.
seq[7, 2, 3, 7, 5, 6, 0, 1]
seq[::2][7, 3, 5, 0]
seq[1::2][2, 7, 6, 1]
We can think of the trailing :2 in the preceding code cells as “count by 2”. Therefore, the 1::2 slice:
- Starts at
1 - Stops at the end because of the first
: - Counts by 2 becauase of the trailing
:2
A clever use of this is to pass -1, which has the useful effect of reversing a list or tuple.
seq[::-1][1, 0, 6, 5, 7, 3, 2, 7]
We will use slicing (subsetting) all semester, so we must understand the examples above.
dict
dict is likely the most important built-in Python data structure. A more common name for it is hash map or associative array. It is a flexibly sized collection of key-value pairs, where key and value are Python objects. One approach for creating one is to use curly braces {} and colons to separate keys and values.
Elements in dictionaries have named keys, while elements in tuples and lists have numerical indices. Dictionaries are handy for passing named arguments and returning named results.
empty_dict = {}
empty_dict{}
A dictionary is a set of key-value pairs.
d1 = {'a': 'some value', 'b': [1, 2, 3, 4]}
d1{'a': 'some value', 'b': [1, 2, 3, 4]}
d1['a']'some value'
d1[7] = 'an integer'
d1{'a': 'some value', 'b': [1, 2, 3, 4], 7: 'an integer'}
We access dictionary values by key names instead of key positions.
You can delete values either using the del keyword or the pop method (which simultaneously returns the value and deletes the key).
d1[5] = 'some value'
d1['dummy'] = 'another value'
d1{'a': 'some value',
'b': [1, 2, 3, 4],
7: 'an integer',
5: 'some value',
'dummy': 'another value'}
del d1[5]
d1{'a': 'some value',
'b': [1, 2, 3, 4],
7: 'an integer',
'dummy': 'another value'}
ret = d1.pop('dummy')ret'another value'
d1{'a': 'some value', 'b': [1, 2, 3, 4], 7: 'an integer'}
The keys and values method give you iterators of the dict’s keys and values, respectively. While the key-value pairs are not in any particular order, these functions output the keys and values in the same order.
d1.keys()dict_keys(['a', 'b', 7])
d1.values()dict_values(['some value', [1, 2, 3, 4], 'an integer'])
List, Set, and Dict Comprehensions
We will focus on list comprehensions, which are Pythonic.
List comprehensions are one of the most-loved Python language features. They allow you to concisely form a new list by filtering the elements of a collection, transforming the elements passing the filter in one concise expression. They take the basic form:
[expr for val in collection if condition]This is equivalent to the following for loop:
result = [] for val in collection: if condition: result.append(expr)The filter condition can be omitted, leaving only the expression.
strings = ['a', 'as', 'bat', 'car', 'dove', 'python']We could use a for loop over strings to keep only strings longer than two and then to capitalize them.
caps = []
for x in strings:
if len(x) > 2:
caps.append(x.upper())
caps['BAT', 'CAR', 'DOVE', 'PYTHON']
A list comprehension is more Pythonic and replaces four lines of code with one. The general format of a list comprehension is [operation on x for x in list if condition]
[x.upper() for x in strings if len(x) > 2]['BAT', 'CAR', 'DOVE', 'PYTHON']
Here is another example. The following code is a for loop and an equivalent list comprehension that squares the integers from 1 to 10.
squares = []
for i in range(1, 11):
squares.append(i ** 2)
squares[1, 4, 9, 16, 25, 36, 49, 64, 81, 100]
[i**2 for i in range(1, 11)][1, 4, 9, 16, 25, 36, 49, 64, 81, 100]
What if we wanted the squares of even numbers?
[i**2 for i in range(1, 11) if i%2==0][4, 16, 36, 64, 100]
[i**2 for i in range(2, 11, 2)][4, 16, 36, 64, 100]
Functions
Functions are the primary and most important method of code organization and reuse in Python. As a rule of thumb, if you anticipate needing to repeat the same or very similar code more than once, it may be worth writing a reusable function. Functions can also help make your code more readable by giving a name to a group of Python statements.
Functions are declared with the def keyword and returned from with the return keyword:
def my_function(x, y, z=1.5): if z > 1: return z * (x + y) else: return z / (x + y)There is no issue with having multiple return statements. If Python reaches the end of a function without encountering a return statement, None is returned automatically.
Each function can have positional arguments and keyword arguments. Keyword arguments are most commonly used to specify default values or optional arguments. In the preceding function, x and y are positional arguments while z is a keyword argument. This means that the function can be called in any of these ways:
my_function(5, 6, z=0.7) my_function(3.14, 7, 3.5) my_function(10, 20)The main restriction on function arguments is that the keyword arguments must follow the positional arguments (if any). You can specify keyword arguments in any order; this frees you from having to remember which order the function arguments were specified in and only what their names are.
Returning Multiple Values
We can write Python functions that return multiple objects. The function f() below returns one tuple that we can unpack to multiple objects.
def f():
a = 5
b = 6
c = 7
return (a, b, c)f()(5, 6, 7)
If we want to return multiple objects with names or labels, we can return a dictionary.
def f():
a = 5
b = 6
c = 7
return {'a' : a, 'b' : b, 'c' : c}f(){'a': 5, 'b': 6, 'c': 7}
f()['a']5
Anonymous (Lambda) Functions
Python has support for so-called anonymous or lambda functions, which are a way of writing functions consisting of a single statement, the result of which is the return value. They are defined with the lambda keyword, which has no meaning other than “we are declaring an anonymous function.”
I usually refer to these as lambda functions in the rest of the book. They are especially convenient in data analysis because, as you’ll see, there are many cases where data transformation functions will take functions as arguments. It’s often less typing (and clearer) to pass a lambda function as opposed to writing a full-out function declaration or even assigning the lambda function to a local variable.
Lambda functions are Pythonic and let us to write simple functions on the fly.
strings = ['foo', 'card', 'bar', 'aaaa', 'abab']strings.sort()
strings['aaaa', 'abab', 'bar', 'card', 'foo']
len(strings[0])4
strings.sort(key=len)
strings['bar', 'foo', 'aaaa', 'abab', 'card']
For example, we could use a lambda function to sort strings by the last letter of each string.
strings.sort(key=lambda x: x[-1])
strings['aaaa', 'abab', 'card', 'foo', 'bar']
What if we want to sort by the second to last letter?
strings.sort(key=lambda x: x[-2])
strings['aaaa', 'abab', 'bar', 'foo', 'card']