Back to teaching portfolio

Chapter 8 · Strings & Text Processing

CSBP119 Practical Questions

String manipulation cleans raw inputs before analytics. These labs emphasize trimming, case folding, searching, and frequency analysis.

Problem 1 · Profile Cleanup

Normalize Student Names

Given a list of names with inconsistent capitalization and whitespace, return a cleaned list of “First Last” strings.

  1. Use list comprehension with strip() and title().
  2. Handle extra internal spaces.
  3. Show the cleaned names.
raw_names = ["  aisha al taj ", "mOHAMMED   omar", "Lina  Saif"]

cleaned = []
for entry in raw_names:
    trimmed = entry.strip()
    words = trimmed.split()
    fixed = ""
    word_index = 0
    while word_index < len(words):
        if word_index > 0:
            fixed += " "
        fixed += words[word_index].capitalize()
        word_index += 1
    cleaned.append(fixed)

print(cleaned)
Problem 2 · Keyword Counts

Survey Keyword Frequency

Ask for free-form text until the user presses Enter on a blank line, then count how often the words “ai”, “data”, and “team” appear.

  1. Store responses in a list.
  2. Lowercase and remove special characters when tallying.
  3. Print a dictionary of keyword counts.
keywords = {"ai": 0, "data": 0, "team": 0}
responses = []

while True:
    line = input("Survey response (blank to stop): ").strip()
    if not line:
        break
    responses.append(line)

remove_chars = ".,!?;:"
for response in responses:
    lowered = response.lower()
    cleaned = ""
    index = 0
    while index < len(lowered):
        char = lowered[index]
        if char not in remove_chars:
            cleaned += char
        index += 1
    words = cleaned.split()
    for word in words:
        if word in keywords:
            keywords[word] += 1

print(keywords)
Problem 3 · Hashtag Scan

Extract Hashtags

Given a block of social posts, extract every hashtag (words that begin with #) and store them in a list for reporting.

  1. Split the text into tokens.
  2. Keep tokens that start with # and have at least two characters.
  3. Print the final list.
posts = "Excited for #hackathon and #UAEU showcase! Share your #project updates."
hashtags = []

tokens = posts.split()
index = 0
while index < len(tokens):
    token = tokens[index]
    if token.startswith("#") and len(token) > 1:
        hashtags.append(token.rstrip(".,!?:;"))
    index += 1

print("Hashtags found:", hashtags)