Create, manipulate, slice, and inspect multi-dimensional arrays (ndarray)
Understand array memory allocation, specifically the distinction between views and copies during slicing
Perform vectorised arithmetic, broadcasting, axis aggregations, and basic matrix operations
Read and write NumPy array data from and to external files
19.1 What Is NumPy?
NumPy (Numerical Python) is a foundational Python library for numerical and scientific computing. It provides efficient data structures and functions for working with large, multi-dimensional numerical datasets. NumPy is widely used in data science, machine learning, engineering, physics, and finance, and it underpins many other libraries in the Python scientific ecosystem, such as pandas, SciPy, and scikit-learn.
You will likely need NumPy in many of your future Python courses, including machine learning and data science, which is why we introduce it here.
19.2 Base Data Structure: The ndarray
The core data structure in NumPy is the ndarray (N-dimensional array).
Key characteristics:
Stores elements of a single data type (homogeneous, e.g., all integers or all floats)
Supports one-dimensional (vectors), two-dimensional (matrices), and higher-dimensional arrays
Optimised for high performance and memory efficiency
19.2.1 Array Creation Examples
There are several ways to create arrays depending on your needs:
arr = np.array([[1, 2, 3], [4, 5, 6]])print(arr.shape) # (2, 3) — 2 rows, 3 columnsprint(arr.ndim) # 2 — number of dimensionsprint(arr.dtype) # int64 — data type of elementsprint(arr.size) # 6 — total number of elements
(2, 3)
2
int64
6
19.3 Array Memory: Views vs. Copies
A crucial concept in NumPy is the difference between a view and a copy when slicing arrays.
When you slice a standard Python list, Python creates a brand-new list containing copies of the elements. However, when you slice a NumPy ndarray, NumPy returns a view of the original array. A view shares the exact same underlying memory block as the original array.
19.3.1 Why Views Matter
Because views share memory, modifying a slice modifies the original array.
# Original arrayoriginal = np.array([10, 20, 30, 40, 50])print(original)# Extract a slice (this is a VIEW)sub_array = original[1:4] # array([20, 30, 40])print("sub_array:")print(sub_array)# Modify an element in the slicesub_array[0] =999# The original array is altered!print("original is altered!")print(original) # Output: array([10, 999, 30, 40, 50])
[10 20 30 40 50]
sub_array:
[20 30 40]
original is altered!
[ 10 999 30 40 50]
This design choice makes NumPy extremely fast and memory-efficient when working with gigabytes of data, as it avoids making unnecessary copies in memory.
19.3.2 Creating an Explicit Copy
If you need an independent array where modifications do not affect the original dataset, you must explicitly create a copy using .copy():
original = np.array([10, 20, 30, 40, 50])# Create an explicit copyindependent_slice = original[1:4].copy()# Modify the copyindependent_slice[0] =999# The original array remains untouchedprint(original) # Output: array([10, 20, 30, 40, 50])print(independent_slice) # Output: array([999, 30, 40])
[10 20 30 40 50]
[999 30 40]
19.4 Common Operations
Element-wise Operations & Vectorised Functions
Unlike standard Python lists, NumPy arrays allow you to perform mathematical operations on all elements simultaneously without writing explicit for loops.
Common methods include reading numerical values from text or CSV files:
# Reading from CSV (skipping header row if present)data = np.loadtxt("data.csv", delimiter=",", skiprows=1)
or reading from NumPy’s native binary format (.npy or .npz):
data = np.load("data.npy")
Saving Data
# Saving to a text/CSV file with formatted outputnp.savetxt("output.csv", data, delimiter=",", fmt="%.2f")# Saving to NumPy's efficient binary formatnp.save("output.npy", data)
NumPy is often combined with pandas, another package designed for more complex data ingestion tasks (such as handling mixed data types, missing values, and column names).
19.6 Advantages and Disadvantages
Advantages
Disadvantages
Extremely fast for numerical operations due to C-based implementation
Limited support for non-numeric or mixed data types
Memory-efficient compared to native Python lists
Less intuitive for labelled or relational data
Extensive mathematical and array processing capabilities
Steeper initial learning curve than basic Python lists
Seamlessly integrates with the scientific Python ecosystem
19.7 Common Alternatives and Their Use Cases
pandas
Used for labelled, tabular data (DataFrames). Builds on top of NumPy and adds indexing, grouping, and data-cleaning tools.
SciPy
Extends NumPy with advanced scientific algorithms (optimisation, signal processing, numerical integration, and statistics).
TensorFlow / PyTorch
Used for large-scale numerical computation and machine learning, especially with GPU acceleration and automatic differentiation.
Python Lists
Suitable for small datasets or heterogeneous (mixed-type) data, but inefficient for large numerical computations.
19.8 Documentation and Resources
The NumPy documentation is very thorough and readable. Practice using the official documentation to look up functions, parameters, and methods.
Create a 5x5 matrix filled with random floating-point values between 0 and 1.
Normalise the matrix so that all values range between 0 and 1 using Min-Max normalisation: \[\text{Normalised} = \frac{X - X_{\text{min}}}{X_{\text{max}} - X_{\text{min}}}\]
HintHint
Consider using np.random.rand(), .min(), and .max().
AnswerAnswer
import numpy as np# Set random seed for reproducibilitynp.random.seed(10)# 1. Create 5x5 random float matrixrand_matrix = np.random.rand(5, 5)print("Original Matrix:\n", rand_matrix)# 2. Min-Max Normalisationmin_val = rand_matrix.min()max_val = rand_matrix.max()normalised_matrix = (rand_matrix - min_val) / (max_val - min_val)print("\nNormalised Matrix:\n", normalised_matrix)