The data model
Quick recap
- Design principles, and the four ways a design rots
- The five principles, at the level of one class
The claim
You have been taught to write classes that are correct. This lecture is about writing classes that are usable.
A language already defines what it means to be a sequence, a number, something you can print, compare, or loop over. Your class can claim those meanings instead of inventing its own.
Agenda
- What the data model is
- Making an object readable:
__repr__and__str__ - Making an object behave like a sequence
- Making an object behave like a number
- The same idea in C++, Java, Rust, Go
- What you take on when you claim a protocol
What the data model is
The set of protocols the interpreter itself uses. A protocol is a name and a promise:
implement __len__, and len() works on your object.
Python's data model covers:
- Objects, values and types
- The standard type hierarchy
- Special methods ← this lecture
- Coroutines
The rule that confuses everyone first
Special methods are called by the interpreter, not by you.
len(deck) # you write this
deck.__len__() # not this
You do not call my_object.__repr__(). You call repr(my_object), and Python dispatches
to the method you wrote.
The dunder is the socket. The built-in function is the plug.
Built-in types vs your types
| numerics | int, float, complex |
| sequences | str, tuple, bytes, list, set, frozenset |
| mappings | dict |
| your types | anything you write with class |
Built-in types get indexing, iteration, arithmetic and printing for free. The data model is how a user-defined type asks for the same treatment.
1. Make it readable
class Lecture:
def __init__(self, instructor, venue, topic):
self.instructor, self.venue, self.topic = instructor, venue, topic
lec = Lecture("Ankush", "CEP-102", "The data model")
print(lec)
<__main__.Lecture object at 0x7d4a1ca31460>
A memory address. Every debugging session with this class starts by losing information.
Two methods, two audiences
class BetterLecture:
def __init__(self, instructor, venue, topic):
self.instructor, self.venue, self.topic = instructor, venue, topic
def __repr__(self):
return f"BetterLecture({self.instructor!r}, {self.venue!r}, {self.topic!r})"
def __str__(self):
return f"{self.topic}, by {self.instructor}"
>>> print(lec)
The data model, by Ankush
>>> lec
BetterLecture('Ankush', 'CEP-102', 'The data model')
>>> [lec]
[BetterLecture('Ankush', 'CEP-102', 'The data model')]
Note !r in the repr. Without it the output is not valid Python, which defeats the
point of a repr you could paste back.
| Called by | For | Should read like | |
|---|---|---|---|
__repr__ |
repr(), the debugger, the REPL, a list's contents |
you | code that could rebuild the object |
__str__ |
print(), str(), f-strings |
the user | a sentence |
If you write only one, write __repr__ — __str__ falls back to it, never the reverse.
2. Make it a sequence
What a list gives you for free:
a = [1, 2, 4, 5, 6, 9]
a[0] # indexing
a[0:3] # slicing
for i in a: ... # iteration
[i*i for i in a] # comprehension
choice(a) # random.choice
shuffle(a) # random.shuffle
sorted(a) # sorting
None of that is special to list. It is the sequence protocol.
The class that asks for none of it
class NotGoodCardDeck:
suits = 'spades diamonds clubs hearts'.split()
ranks = [str(n) for n in range(2, 11)] + list('JQKA')
def __init__(self):
self._cards = [Card(rank, suit) for suit in self.suits for rank in self.ranks]
def shuffle(self): ...
def pick_random_card(self): ...
def pick_cards_by_suit(self, suit): ...
Three methods you now have to write, name, document and test — and a user has to learn.
shuffle already exists. pick_random_card is random.choice.
Every method here reinvents something the language already had a name for.
Two methods instead
Card = collections.namedtuple('Card', ['rank', 'suit'])
class CardDeck:
suits = 'spades diamonds clubs hearts'.split()
ranks = [str(n) for n in range(2, 11)] + list('JQKA')
def __init__(self):
self._cards = [Card(rank, suit) for suit in self.suits for rank in self.ranks]
def __len__(self):
return len(self._cards)
def __getitem__(self, index):
return self._cards[index]
What those two bought
deck = CardDeck()
len(deck) # 52
deck[0] # Card(rank='2', suit='spades')
deck[0:3] # [Card('2','spades'), Card('3','spades'), Card('4','spades')]
Card('Q', 'hearts') in deck # True
Card('Q', 'beasts') in deck # False
next(iter(reversed(deck))) # Card(rank='A', suit='hearts')
choice(deck) # Card(rank='9', suit='diamonds')
sorted(deck, key=spades_high)[-1] # Card(rank='A', suit='spades')
Slicing, membership, iteration, reversed, random.choice, sorted — none of which
were mentioned in the class.
__getitem__ alone makes an object iterable: Python falls back to calling it with
0, 1, 2 … until IndexError.
Reading is not writing
shuffle(deck)
This needs a third method. random.shuffle works in place, so it has to assign:
def __setitem__(self, position, card):
self._cards[position] = card
Without it:
TypeError: 'CardDeck' object does not support item assignment
Protocols are graded. Read-only sequence is one contract; mutable sequence is a larger one. Claim only what you can honour.
3. Make it a number
class Vector:
def __init__(self, x, y):
self.x, self.y = x, y
def __str__(self): return f"{self.x}i + {self.y}j"
def __repr__(self): return f"Vector({self.x}, {self.y})"
def __add__(self, other):
return Vector(self.x + other.x, self.y + other.y)
def __mul__(self, scalar):
return Vector(scalar * self.x, scalar * self.y)
def __rmul__(self, scalar): # for `5 * v`, not `v * 5`
return Vector(scalar * self.x, scalar * self.y)
def __abs__(self): return hypot(self.x, self.y)
def __bool__(self): return bool(abs(self))
The result
a, b = Vector(3, 4), Vector(5, 6)
print(a) # 3i + 4j
repr(a) # 'Vector(3, 4)'
a + b # 8i + 10j
a * 5 # 15i + 20j
5 * a # 15i + 20j
abs(a) # 5.0
bool(a) # True
bool(Vector(0, 0)) # False
Two details worth noticing:
__rmul__exists because5 * aasksintfirst.intdoes not know aboutVector, returnsNotImplemented, and Python then asks the right operand. Without__rmul__,a * 5works and5 * araises.__bool__defined as "non-zero magnitude" meansif vector:reads as "is this vector non-zero" — a domain statement, written in language syntax.
A library you already use
spaCy's Doc is a sequence of tokens, and nothing more exotic:
doc = nlp("The iceberg is called the Python data model...")
doc[2] # indexing -> is
doc[5:8] # slicing -> Python data model
[t for t in doc if t.pos_ == "VERB"] # comprehension
for token in reversed(doc): ... # reversed
choice(doc) # random.choice
(Output as recorded in the source notebook.)
There is no doc.get_token_at(2) and no doc.get_slice(5, 8). The API is small because
the language already had names for most of what it does.
In your own code
Three of these are already in your repositories. 22 of your 36 repos overload at least one operator, so this is not a foreign idea — it is one you use halfway.
Equality, written out four times
Param2725/snake_game
declares the type and stops there:
struct Position { int x, y; }; // no operator==
So every comparison is spelled out — snake.cpp:244-249:
for (auto &o : obstacles) if (newHead.x == o.x && newHead.y == o.y) gameOver = true;
for (auto &s : snake) if (newHead.x == s.x && newHead.y == s.y) gameOver = true;
if (newHead.x == food.x && newHead.y == food.y) { score++; ... }
Four coordinate comparisons in six lines. Add a z and every one of them is a bug.
One of you already fixed it
struct Position {
int x, y;
Position(int x = 0, int y = 0) : x(x), y(y) {}
bool operator==(const Position& o) const { return x == o.x && y == o.y; }
bool operator!=(const Position& o) const { return !(*this == o); }
};
Four lines, written once. The three checks above become:
newHead == o newHead == s newHead == food
operator== is C++'s __eq__. The comparison did not get shorter by accident — it got
shorter because the type now knows what equality means.
Halfway there is still by hand
Same repo,
snakeCycle.cpp:326-333.
They defined operator== and still wrote the search:
bool checkSelfCollision() const {
if (body.size() <= 1) return false;
Position head = body[0];
for (size_t i = 1; i < body.size(); i++) {
if (head == body[i]) return true;
}
return false;
}
The standard library already knows this loop:
bool checkSelfCollision() const {
return std::find(body.begin() + 1, body.end(), body.front()) != body.end();
}
Both agree on every input. Eight lines describing how, replaced by one line saying
what — and it works only because operator== was already there.
In Python the same rule is head in body[1:], via __contains__.
The identical loop appears in five repos: dudhatmonar:65, Heer117:65,
LoveShah21:409, kkhushie:173, CharmiBhayani:326.
A method that hides a method
Eight repos write some version of this:
int getLength() const { return body.size(); } // CharmiBhayani:335, maahirgit:138
int length() const { return body.size(); } // dudhatmonar:79
size_t getLength() const { return body.size(); } // maitry4:245
The whole body is a call to the thing it is hiding, under a name every reader has to learn — and a different name in each repo.
size() on the class, or __len__ in Python, and len(snake) works for everyone with
nothing to learn.
💬 Your snake game, as a sequence
The snake body is a list of coordinates. You wrote methods to grow it, check collisions, and draw it.
Which dunder methods would let you delete code you already wrote?
Answer
| Method | Deletes |
|---|---|
__len__ |
getLength(), size(), bodyCount() |
__getitem__ |
getSegment(i), head() becomes body[0] |
__contains__ |
the self-collision loop becomes if head in body[1:] |
__iter__ |
every for i in range(len(body)) in the renderer |
The self-collision check is the striking one. Most implementations write a loop
comparing the head against each segment. With __contains__, the rule becomes one
line that says what it means rather than how it is computed.
A caution: only claim __getitem__ if indexing your object genuinely makes sense.
A protocol you implement badly is worse than one you never claimed, because now
every built-in that trusts it is wrong too.
The same idea elsewhere
Python's version is unusually broad, but the idea is not Python's.
| Language | Mechanism | Examples |
|---|---|---|
| C++ | operator overloading, iterators | operator+, operator[], operator<<, begin()/end() |
| Java | interfaces | toString, equals/hashCode, Comparable, Iterable, AutoCloseable |
| C# | operators + interfaces | operator+, indexers this[int i], IEnumerable, IDisposable |
| Rust | traits | Display, Debug, Add, Index, Iterator, PartialOrd |
| Go | implicit interfaces | Stringer, sort.Interface, io.Reader/io.Writer |
The same class, in C++
class Vector {
public:
Vector operator+(const Vector& o) const { return {x + o.x, y + o.y}; }
friend Vector operator*(double s, const Vector& v) { return {s*v.x, s*v.y}; }
friend std::ostream& operator<<(std::ostream& os, const Vector& v) {
return os << v.x << "i + " << v.y << "j";
}
private:
double x, y;
};
operator<< is __str__. The free operator* taking the scalar first is __rmul__.
Different spelling, identical idea: the type declares that it participates in an
existing vocabulary.
Go states the principle most plainly — a type implements Stringer merely by having a
String() string method. No inheritance, no declaration.
What you take on
Claiming a protocol is a promise to everything that already uses it.
| You implement | You promise |
|---|---|
__eq__ |
equal objects stay equal, and hash the same |
__lt__ |
a consistent ordering, or sorted misbehaves |
__len__ |
a non-negative integer, cheap to compute |
__getitem__ |
IndexError past the end, or iteration never stops |
A built-in function does not verify your promise. It relies on it.
That is the trade the whole lecture rests on: you get an enormous amount of existing code for free, and in exchange you are held to the contract that code was written against.
Closing
import this
Beautiful is better than ugly. Simple is better than complex. Special cases aren't special enough to break the rules.
Three questions to ask of any class you write from here:
- Does it print as something a human can read?
- Does it reinvent a verb the language already has?
- If it behaves like a sequence or a number, does it say so?
References:
- Chapter 1, The Python Data Model — Fluent Python, Luciano Ramalho
- Python reference — Data model
- Difference between
__str__and__repr__ - spaCy v3: design concepts explained