Stringing Together

Fred Agbo

September 18, 2026

Happy Friday! Quick Announcements

  • We will continue with class group activities on Monday!
  • Problem sets 1 and 2 are all graded.
    • impressive performance overall!
      • Many did not fill out the meta data portion on each file
      • Many also did not provide good documentation (in-code and when running journal)
    • Check and monitor your grades on Canvas.
  • Problem set 3 is due next week Monday at 10 pm
  • Remember the only way to master coding is to keep practicing

Floating Representations

  • Python represents floating point (fractional) numbers using two integers
    • One to represent the significant digits
    • One to represent the exponent (where the decimal place is)
  • \(1\frac{1}{4}\) Example
    • In decimal: \(\quad\displaystyle 1\frac{1}{4} = \frac{1}{1} + \frac{2}{10} + \frac{5}{100} = 1.25 = (125, -2)\)
    • In binary: \(\quad\displaystyle 1\frac{1}{4} = \frac{1}{1} + \frac{0}{2} + \frac{1}{4} = 1.01 = (101, -10)\)

Problem 3: Floating Problems

  • With this in mind, how could you convert the value \(\tfrac{7}{8}\) to a binary floating point representation? \[\frac{7}{8} = \frac{0}{1} + \frac{1}{2} + \frac{1}{4} + \frac{1}{8} = 0.111 = (111, -11)\]
  • Now how would we convert \(\frac{1}{10}\) to binary??
    • We run into a problem! An infinitely repeating sequence! \[\frac{1}{10} = \frac{0}{1} + \frac{0}{2} + \frac{0}{4} + \frac{0}{8} + \frac{1}{16} + \frac{1}{32} + \frac{0}{64} + \frac{0}{128} + \frac{1}{256} + \cdots = 0.0001100110011\ldots\]
    • Have to stop the sequence somewhere and approximate it: \[\frac{3}{32} = 0.09375\quad\text{or}\quad\frac{25}{256} = 0.09765625\]

Consequences

  • The best we can do within the range of normal integers \[\frac{3602879701896397}{2^{55}} = 0.10000000000000000555111512312578270\]
  • When doing operations on these numbers, extra decimals will sometimes get rounded off, suddenly making the number look precise, but you might always have a tiny bit of this rounding error showing up in floating point values.
  • So be careful using == for floating numeric comparisons! Rounding might result in unexpected falsehoods
    • 0.1 + 0.1 + 0.1 != 0.3
    • Far better to check if two numbers are within a small margin of one another, or greater or less than the other

Other Data Types (STRING)

  • Numbers are great, but what about other types of data?
  • We will introduce STRING as an a data type
  • Note! other data type you would see later include:
    • lists!

Strings

  • A string in Python represents textual data, in form of a sequence of individual characters
    • Domain: all possible sequences of characters
    • Operations: Many! We’ll see some of them soon
  • Denoted by placing the desired sequence of characters between two quotation marks
    • 'I am a string'
    • In Python, either single or double quotes can be used, but the ends must match
      • "I am also a string!"
      • "I'm sad you've gone"

Sequences

  • Both strings and lists are examples of a more general type called a sequence
    • Strings are sequences of characters
    • Lists are sequences of anything
  • Sequences are ordered, so we can number off their elements, which we call their index
    • Counting in Python always starts with 0, so the first element of the sequence has index 0
  • Python defines operations that work on all sequences
    • Selecting an individual element out of a sequence
    • Concatenating two sequences together
    • Determing the number of elements in a sequence

Selection

  • You can select or “pluck out” just a single element from a sequence using square brackets [ ]
    • There are no commas between these square brackets, so they can’t be confused with a list
    • The square brackets come after the sequence (or variable name representing a sequence)
    • Inside the square brackets, you place the index number of the element you want to select
>>> A = [2, 4, 6, 8]
>>> print(A[1])
4
>>> B = "Spaghetti"
>>> print(B[6])
't'

Concatenation

  • Concatenation is the act of taking two separate objects and bringing them together to create a single object
  • For sequences, concatenation takes the contents of one sequence and add them to the end of another sequence
  • The + operator concatenates sequences
    • This is why it is important to keep track of your variable types! + will add two integers, but will concatenate two strings
>>> 'fish' + 'sticks'
'fishsticks'
>>> A = [1, 'fish']
>>> B = [2, 'fish']
>>> print(A + B)
[1, 'fish', 2, 'fish']

Lengths

  • The number of elements in a sequence is commonly called its length, and can be given by the len( ) function

  • Simply place the sequence you desire to know the length of between the parentheses:

    >>> len("spaghetti")
    9
  • You can have sequences of 0 length as well!

    >>> A = ""
    >>> B = [ ]
    >>> print( len(A) + len(B) )
    0

Representing Characters

  • We use numeric encodings to represent character data inside the machine, where each character is assigned an integer value.
  • Character codes are not very useful unless standardized though!
    • Competing encodings in the early years made it difficult to share data across machines
  • First widely adopted character encoding was ASCII (American Standard Code for Information Interchange)
  • Originally just with 128 possible characters, even after expanding to 256, ASCII proved inadequate in the international world, and has therefore been superseded by Unicode.

ASCII

image/svg+xml 0 1 2 3 4 5 6 7 8 9 A B C D E F 0x 1x 2x 3x 4x 5x 6x 7x \0 \b \t \n \v \f \r ! " # $ % & ' ( ) * + , - . / 0 1 2 3 4 5 6 7 8 9 : ; < = > ? @ A B C D E F G H I J K L M N O P Q R S T U V W X Y Z [ \ ] ^ _ ` a b c d e f g h i j k l m n o p q r s t u v w x y z { | } ~
The ASCII subset of Unicode

Meeting chr and ord

  • Python includes two build in functions to simplify conversion between an integer and the corresponding Unicode character
  • chr takes a base-10 integer and returns the corresponding Unicode character as a string
    • chr(65) gives "A" (capital A)
    • chr(960) gives "π" (Greek letter pi)
  • ord goes the other direction, taking a single character string and returning the corresponding base-10 integer of that character in Unicode
    • ord("B") gives 66
    • ord(" ") gives 32
    • ord("π") gives 960

Live-Coding 1

Concatenating ASCII

  • Suppose we wanted to print out the simple ASCII table as we saw it in the video and above slide
  • Grid of 8 rows and 16 columns
  • We can use chr to get the characters
  • Need to “build-up” a string to print for each row

A Solution

row=''
for i in range(32,127):
    if len(row) == 16:
        print(row)
        row = ""
    row = row + chr(i)
print(row)

Abstract Strings

  • Characters (and their Unicode representation) are most often used in programming when combined to make collections of consecutive characters called strings.
  • Internally, strings are stored as a sequence of characters in a sequential chunk of memory.
  • You don’t have to (and generally don’t want to) think of the internal representation.
    • Better to think of the string as a single abstract unit
  • Python emphasizes this abstract view by defining a built-in string class that already defines a selection of higher-level operations on string objects

Character Picking Recap

  • A string is an ordered collection of characters
    • Character positions in the string are identified by an index, which starts at 0

  • You can select individual characters from the string using the syntax

    string[k]

    where string is the variable assigned to the desired string and k is the index integer of the character you want

    >>> print("spaghetti sauce"[5])
    e

Back it Up

  • Sometimes it is more useful to count from the end of the string, not the beginning
  • Python gives you a convenient way to do this, using negative indexes


  • A common use case is to grab the last character of the string, using

    s[-1]

    which is shorthand for

    s[len(s)-1]

Slicing

  • Often, you may want more than a single character

  • Python allows you to specify a starting and an ending index through an operation known as slicing

  • The syntax looks like:

    string_variable[start : limit]

    where start is the first index to be included and everything up to but not including the limit is included

  • start and limit are actually optional (but the : is not)

    • If start omitted, the slice will begin at the start of the string
    • If limit omitted, the slice will proceed to the end of the string

and Dicing

  • Can add a third component to the slice syntax, called a stride

    string_variable[start : limit : stride]
  • Specifies how large the steps are between each included index

  • Can also make the stride negative to proceed backwards through a string

    >>> s = "spaghetti sauce"
    >>> s[4:8]
    hett
    >>> s[10:]
    sauce
    >>> s[:10:2]
    sahti

Repeat again?

  • We’ve already seen how we can use addition (+) in Python to concatenate strings
  • In math, adding something many times is the same as multiplying

\[5+5+5+5+5+5 = 6 \times 5\]

  • The same logic holds true for Python strings!
    • You multiply by a integer: the number of times you want the concatenation repeated
    • You can not multiply two strings together, Python will not understand what you are trying to do
    print("Betelguese, " * 3)

Comparing Strings

  • Python lets you use normal comparison operators to compare strings

    string1 == string2

    is true if string1 and string2 contain the same characters in the same order

  • Comparisons involving greater than or less than are done similar to alphabetical ordering

    • Start at the beginning and compare a character. If they are the same, then compare the next character, etc
  • All comparisons are done according to their Unicode values.

    • Called lexicographic ordering
    • "cat" > "CAT"

Can’t change a string’s colors

  • Strings are what we call immutable: they can not be modified in place by clients.

  • You can “look” at different parts of the string, but you can not “change” those parts without making a whole new string

    s = "Cats!"
    s[0] = "R"   # THIS WILL ERROR!!
  • You can of course create a new string object with the desired traits:

    s = "R" + s[1:]
  • This applies to all methods that act on strings as well: they return a new string, they do not modify the original

Methods to find string patterns

Method Description
string.find(pattern) Returns the first index of pattern in string, or -1 if it does not appear
string.find(pattern, k) Same as the one-argument version, but starts searching at index k
string.rfind(pattern) Returns the last index of pattern is string, or -1 if missing
string.rfind(pattern, k) Same as the one-argument version, but searches backwards from index k
string.startswith(prefix) Returns True if the string starts with prefix
string.endswith(suffix) Returns True if the string ends with suffix

Transforming Methods

Method Description
string.lower() Returns a copy of string with all letters converted to lowercase
string.upper() Returns a copy of string with all letters converted to uppercase
string.capitalize() Returns a copy of string with the first character capitalized and the rest lowercase
string.strip() Returns a copy of string with whitespace and non-printing characters removed from both ends
string.replace(old, new) Returns a copy of string with all instances of old replaced by new

Classifying Character Methods

Method Description
char.isalpha() Returns True if char is a letter
char.isdigit() Returns True if char is a digit
char.isalnum() Returns True if char is letter or a digit
char.islower() Returns True if char is a lowercase letter
char.isupper() Returns True if char is an uppercase letter
char.isspace() Returns True if char is a whitespace character (space, tab, or newline)
char.isidentifier() Returns True if char is a legal Python identifier

Live-Coding 2

Igpay Atinlay

  • Suppose we wanted to write a script that converted English to Pig Latin
  • Rules of Pig Latin:
    • If the word begins with a consonant, move everything up to the first vowel to the end and append on “ay” at the end
      fleet ⟶ eetflay
    • If the word starts with a vowel, just append “way” to the end
      orange ⟶ orangeway
    • If the word has no vowels, do nothing
  • Bonus: What English words, when “Pig-Latin-ified” are still valid English words?

A Solution

from english import ENGLISH_WORDS, is_english_word

def find_first_vowel(word):
    """ 
    Finds the first vowel index in a string

    Algorithm:
        Loop through all letters
        Check if that letter is in aeiou
        If it is, immediately return the index
    """
    for i in range(len(word)):
        if word[i].lower() in "aeiou":
            return i
    return -1

def pig_latin(word):
    """
    Converts a word to its Pig Latin form

    Algorithm:
        Check the first letter to see if vowel
            Task 1 if is not vowel
                Figure out where first vowel is
                Slice and rearrange
            Task 2
                Concatenate on a way
    """
    first_vowel = find_first_vowel(word)
    if first_vowel == 0:
        # Easy tack on way
        word += "way"
    elif first_vowel > 0:
        # Find first vowel and rearrange
        first_part = word[:first_vowel]
        second_part = word[first_vowel:]
        word = second_part + first_part + "ay"
    return word


if __name__ == '__main__':
    for word in ENGLISH_WORDS:
        platin = pig_latin(word)
        if is_english_word(platin) and word != platin:
            print(word, platin)

Learning English

  • When working with sequences of characters, it is often useful or desirable to determine if they form actual valid English words
  • This class provides for you a new library, through the file english.py
  • The english library provides two objects you can import into your programs:
    • The constant ENGLISH_WORDS, which is a list of all the valid words in the English dictionary
    • The function is_english_word(), which accepts a single string as an argument and returns True if the string represents a valid English word.
  • This library will be particularly useful for Wordle!

Example: How many 2 letter words?

  • Before we start writing code, let’s pause. Give a physical English dictionary, how could you go about figuring out the number of two letter words?
from english import ENGLISH_WORDS

count = 0
for word in ENGLISH_WORDS:
    if len(word) == 2:
        count += 1
print(count)
from english import is_english_word

count = 0
alphabet = 'abcdefghijklmnopqrstuvwxyz'
for letter1 in alphabet:
    for letter2 in alphabet:
        word = letter1 + letter2
        if is_english_word(word):
            count += 1
print(count)
// reveal.js plugins