Expressing Tensorflow 2.x operations in plain Mathematical Notation
Despite the vast documentation on Tensorflow operations, those are little to no documented in mathematical notation. This makes it hard to abstract and generalize some of those ideas, as well as make it rather impossible to make those operations accessible for the common mathematician. Here is a little subset of these operations documented in such a fashion. Furthermore, I also note how you can actually use Tensor Algebra to express theta-join operations. You can find a full list of Tensorflow operations here.
1. Element-wise Transformations
These operations apply a scalar function $f: \mathbb{R} \to \mathbb{R}$ independently to each component of a tensor without changing its dimensions.
Mathematical Definition
Given an $N$-order tensor $\mathcal{X} \in \mathbb{R}^{I_1 \times I_2 \times \dots \times I_N}$, the transformation $\mathcal{Y} = f(\mathcal{X})$ is defined component-wise as:
\[\mathcal{Y}_{i_1 i_2 \dots i_N} = f\left(\mathcal{X}_{i_1 i_2 \dots i_N}\right)\]where $i_k \in {1, 2, \dots, I_k}$ for each mode $k$.
TensorFlow Implementation
import tensorflow as tf
# Example using a sigmoid activation function
Y = tf.math.sigmoid(X)
2. Tensor Contractions & Matrix Multiplications
These operations compute inner products over specified matching dimensions, effectively reducing the rank of the combined structures.
Mathematical Definition
General Tensor Contraction
If we contract an order-$P$ tensor $\mathcal{A}$ and an order-$Q$ tensor $\mathcal{B}$ over their last $k$ and first $k$ axes respectively, using Einstein summation convention (implicit summation over repeated indices):
\[\mathcal{C}_{i_1 \dots i_{P-k} j_{k+1} \dots j_Q} = \mathcal{A}_{i_1 \dots i_{P-k} m_1 \dots m_k} \mathcal{B}^{m_1 \dots m_k}_{\phantom{m_1 \dots m_k} j_{k+1} \dots j_Q}\]Batch Matrix Multiplication
For two 3D tensors representing batches of matrices, $\mathcal{A} \in \mathbb{R}^{B \times I \times J}$ and $\mathcal{B} \in \mathbb{R}^{B \times J \times K}$:
\[\mathcal{C}_{b i k} = \sum_{j=1}^{J} \mathcal{A}_{b i j} \mathcal{B}_{b j k} \implies \mathcal{C}_{b i k} = \mathcal{A}_{b i j} \mathcal{B}_{b \phantom{j} k}^{\phantom{b} j}\]TensorFlow Implementation
# General Contraction (contracting over last 2 axes of A and first 2 axes of B)
C_dot = tf.tensordot(A, B, axes=2)
# Batch Matrix Multiplication
C_matmul = tf.linalg.matmul(A, B)
3. Axis Permutations and Transpositions
These transformations map tensor coordinates via a permutation function $\pi$, rearranging the axes layout.
Mathematical Definition
Given $\mathcal{X} \in \mathbb{R}^{I_1 \times I_2 \times \dots \times I_N}$ and a bijection $\pi$ of the set of indices ${1, 2, \dots, N}$, the transformed tensor $\mathcal{Y} = \text{permute}(\mathcal{X}, \pi)$ is defined by:
\[\mathcal{Y}_{i_{\pi(1)} i_{\pi(2)} \dots i_{\pi(N)}} = \mathcal{X}_{i_1 i_2 \dots i_N}\]TensorFlow Implementation
# Permuting axes mapping index positions [0, 1, 2] -> [2, 0, 1]
Y = tf.transpose(X, perm=[2, 0, 1])
4. Tensor Reductions
Reduction transformations collapse specific dimensions by aggregating elements via a binary operator $\bigoplus$ (such as $\sum, \max, \prod$).
Mathematical Definition
If reducing $\mathcal{X} \in \mathbb{R}^{I_1 \times I_2 \times I_3}$ along the second axis ($I_2$) using summation:
\[\mathcal{Y}_{i_1 i_3} = \sum_{i_2=1}^{I_2} \mathcal{X}_{i_1 i_2 i_3} = \mathcal{X}_{i_1 i_2 i_3} \mathbf{1}^{i_2}\]TensorFlow Implementation
# Reduction along axis 1
Y = tf.reduce_sum(X, axis=1)
5. Linear Convolution Transforms
Spatial transforms that slide a localized kernel filter tensor across a multi-dimensional data tensor.
Mathematical Definition
For a 4D input $\mathcal{X}$ (Batch $b$, Height $h$, Width $w$, Channels $c$) and a filter kernel $\mathcal{K}$ (Filter Height $k_h$, Filter Width $k_w$, Input Channels $c$, Output Filters $f$), with strides $S_h, S_w$:
\[\mathcal{Y}_{b, \, h, \, w, \, f} = \sum_{\delta h} \sum_{\delta w} \sum_{c} \mathcal{X}_{b, \, h \cdot S_h + \delta h, \, w \cdot S_w + \delta w, \, c} \cdot \mathcal{K}_{\delta h, \, \delta w, \, c, \, f}\]TensorFlow Implementation
Y = tf.nn.conv2d(X, K, strides=[1, 1, 1, 1], padding='SAME')
6. Tensor Slicing
Slicing extracts contiguous sub-tensors by selecting bounded index intervals along target dimensions.
Mathematical Definition
Given $\mathcal{X} \in \mathbb{R}^{I_1 \times I_2 \times \dots \times I_N}$, a slice with starting indices $s_k$, ending indices $e_k$, and steps $t_k$ creates an output $\mathcal{Y} \in \mathbb{R}^{J_1 \times J_2 \dots \times J_N}$:
\[\mathcal{Y}_{j_1 j_2 \dots j_N} = \mathcal{X}_{(s_1 + j_1 \cdot t_1)(s_2 + j_2 \cdot t_2)\dots(s_N + j_N \cdot t_N)}\]where $j_k \in {0, 1, \dots, \lfloor \frac{e_k - s_k - 1}{t_k} \rfloor}$.
TensorFlow Implementation
# Slicing a sub-tensor using standard Python slicing notation
Y = X[s_1:e_1:t_1, s_2:e_2:t_2]
7. Tensor Squeezing
Squeezing eliminates singleton dimensions (axes of size 1) without mutating internal value order.
Mathematical Definition
Let $\mathcal{X} \in \mathbb{R}^{I_1 \times \dots \times I_N}$ where a subset of axes $A$ satisfies $I_a = 1, \forall a \in A$. If $\phi: {1, \dots, M} \to {1, \dots, N} \setminus A$ is a strictly increasing monotonic index mapping:
\[\mathcal{Y}_{j_1 j_2 \dots j_M} = \mathcal{X}_{i_1 i_2 \dots i_N} \quad \text{where } i_k = \begin{cases} 1 & \text{if } k \in A \\ j_{\phi^{-1}(k)} & \text{if } k \notin A \end{cases}\]TensorFlow Implementation
# Removes all dimensions of size 1
Y = tf.squeeze(X, axis=list(A))
8. Tensor Aggregation
Advanced grouping transformations over segmented regions or cluster mappings.
Mathematical Definition
For a matrix $\mathcal{X} \in \mathbb{R}^{I \times J}$ and a segment assignment vector $S \in \mathbb{N}^I$, the segmented aggregation mapping to $\mathcal{Y} \in \mathbb{R}^{K \times J}$ evaluates via Kronecker deltas:
\[\mathcal{Y}_{k, j} = \sum_{i=1}^{I} \mathcal{X}_{i, j} \cdot \delta_{k, S_i}\]TensorFlow Implementation
Y = tf.math.segment_sum(X, segment_ids=S)
9. One-Hot Encoding
Expands a discrete label tensor by adding a categorical indicator axis of depth $D$.
Mathematical Definition
Given $\mathcal{X} \in \mathbb{N}^{I_1 \times \dots \times I_N}$, active value $\alpha$, and inactive value $\beta$, the expanded encoding $\mathcal{Y} \in \mathbb{R}^{I_1 \times \dots \times I_N \times D}$ along target index $d \in {0, \dots, D-1}$ is:
\[\mathcal{Y}_{i_1 i_2 \dots i_N d} = \alpha \cdot \delta_{d, \, \mathcal{X}_{i_1 i_2 \dots i_N}} + \beta \cdot \left(1 - \delta_{d, \, \mathcal{X}_{i_1 i_2 \dots i_N}}\right)\]TensorFlow Implementation
Y = tf.one_hot(X, depth=D, on_value=alpha, off_value=beta)
10. Tensor Gathering and Scattering
Indexed read and write operations mapping sparse index topologies to dense configurations.
Mathematical Definition
Multi-Dimensional Gather (tf.gather_nd)
Extracts indices from source $\mathcal{X} \in \mathbb{R}^{I_0 \dots \times I_{N-1}}$ using coordinate lookup map $\mathcal{R} \in \mathbb{N}^{J_0 \dots \times J_{M-1} \times K}$:
\[\mathcal{Y}_{j_0 \dots j_{M-1} \, i_K \dots i_{N-1}} = \mathcal{X}_{\mathcal{R}_{j_0 \dots j_{M-1} 0}, \, \dots, \, \mathcal{R}_{j_0 \dots j_{M-1} (K-1)}, \, i_K, \, \dots, \, i_{N-1}}\]Multi-Dimensional Scatter Update (tf.tensor_scatter_nd_add)
Accumulates updates $\mathcal{U}$ into base tensor $\mathcal{X}$ along target indices $\mathcal{R}$:
\[\mathcal{Y}_{i_0 \dots i_{N-1}} = \mathcal{X}_{i_0 \dots i_{N-1}} + \sum_{j_0 \dots j_{M-1}} \mathcal{U}_{j_0 \dots j_{M-1} \, i_K \dots i_{N-1}} \cdot \prod_{k=0}^{K-1} \delta_{i_k, \, \mathcal{R}_{j_0 \dots j_{M-1} k}}\]TensorFlow Implementation
# Gather multi-dimensional slices
Y_gather = tf.gather_nd(X, R)
# Scatter multi-dimensional updates into an existing tensor
Y_scatter = tf.tensor_scatter_nd_add(X, R, U)
11. Relational $\theta$-Joins Over Tensors
Mimics relational database conditional evaluation across two separate tensors while combining their coordinate spaces.
Mathematical Definition
Given $\mathcal{A} \in \mathbb{R}^{I_1 \times \dots \times I_L}$ with join axis index $\mu$, and $\mathcal{B} \in \mathbb{R}^{J_1 \times \dots \times J_R}$ with join axis index $\nu$:
-
Broadcast Predicate Matrix Calculation ($\mathcal{M}$): \(\mathcal{M}_{i_1 \dots i_L j_1 \dots j_R} = \mathbb{I}\Big(\mathcal{A}_{i_1 \dots i_mu \dots i_L} \,\, \theta \,\, \mathcal{B}_{j_1 \dots j_\nu \dots j_R}\Big)\)
-
Coordinate Extraction via Index Lookup Matrix ($\mathcal{R}$): \(\mathcal{R} = \text{positions}(\mathcal{M} == 1) \in \mathbb{N}^{K \times (L+R)}\)
-
Value Materialization: \(\mathcal{Y}_{k, 0} = \mathcal{A}_{\mathcal{R}_{k, 1}, \dots, \mathcal{R}_{k, L}}, \quad \mathcal{Y}_{k, 1} = \mathcal{B}_{\mathcal{R}_{k, L+1}, \dots, \mathcal{R}_{k, L+R}}\)
TensorFlow Implementation
# Step 1: Reshape to broadcast comparison over targeted axes
# Assumes A is 2D [I, M_axis] and B is 2D [J, N_axis]
A_expanded = tf.expand_dims(tf.expand_dims(A, 1), 2) # Shape: [I, 1, 1, M_axis]
B_expanded = tf.expand_dims(tf.expand_dims(B, 0), 0) # Shape: [1, 1, J, N_axis]
# Compute boolean predicate indicator mask (e.g., theta is strict equality)
M_mask = tf.reduce_all(tf.equal(A_expanded, B_expanded), axis=-1)
# Step 2 & 3: Find valid matching coordinates and materialize entries
R_coords = tf.where(M_mask)
vals_A = tf.gather_nd(A, R_coords[:, :2])
vals_B = tf.gather_nd(B, R_coords[:, 2:])
Y_join = tf.concat([vals_A, vals_B], axis=-1)
12. Complex Contextual Multi-Axis Tensor Expressions
This section demonstrates how to formulate complex, arbitrary multi-axis expressions combining generalized Einstein contractions, cell-wise logical operators ($\vee$), and existential quantifiers (∃) acting over explicit sets of axes.
We formalize the evaluation of a target indicator expression: \(\mathbb{I}(\exists I. \theta(\sum_J M[J;I]))\) where I and J are distinct index tuples (sets of axes) rather than singular dimensions.
Mathematical Definition
Given a tensor $\mathcal{M}$ spanning multi-index layout spaces $\mathbf{j} = (j_1, \dots, j_m) \in J$ and $\mathbf{i} = (i_1, \dots, i_n) \in I$:
-
Multi-Index Contraction over Set J: \(\mathcal{Z}_{\mathbf{i}} = \sum_{j_1} \dots \sum_{j_m} \mathcal{M}_{(j_1, \dots, j_m, \, i_1, \dots, i_n)}\)
-
Existential Quantifier Disjunction over Set I via Predicate θ: \(y = \mathbb{I}\left(\exists \mathbf{i} \in I . \, \theta(\mathcal{Z}_{\mathbf{i}})\right) = \bigvee_{i_1} \dots \bigvee_{i_n} \mathbb{I}\big(\theta(\mathcal{Z}_{(i_1, \dots, i_n)})\big)\)
Alternatively, using an algebraic bounding framework without explicit boolean branching logic: \(y = \text{clip}\left( \sum_{i_1} \dots \sum_{i_n} \mathbb{I}\big(\theta(\mathcal{Z}_{\mathbf{i}})\big), \,\, 0, \,\, 1 \right)\)
For cell-wise logical OR ($\vee$) statements matching conditions such as $\mathcal{A}{ijk} > 0 \vee \mathcal{B}{ijk} > 0$, the indicator evaluates as: \(\text{Logical\_OR}(\mathcal{A}, \mathcal{B}) = \max\Big( \mathbb{I}(\mathcal{A} > 0), \, \mathbb{I}(\mathcal{B} > 0) \Big)\)
TensorFlow Implementation
import tensorflow as tf
# Define numerical axis positions corresponding to index sets J and I
# Example assuming M is a 5D tensor: J maps to axes, I maps to axes [2, 3, 4]
axes_J = [0, 1]
axes_I = [2, 3, 4]
# 1. Compute multi-axis contraction over the designated axis indices in set J
Z = tf.reduce_sum(M, axis=axes_J)
# 2. Evaluate cell-wise condition theta (e.g., matching a numerical boundary)
# Includes element-wise logical OR implementation logic example via '|'
condition_mask = (Z > 0.0) | (Z < -1.0)
# 3. Handle Existential Quantifier ($\exists$ I) collapsing multi-axis set I
exists_condition = tf.reduce_any(condition_mask, axis=axes_I)
# 4. Final step conversion back to a numeric indicator configuration
y_final = tf.cast(exists_condition, dtype=tf.float32)
Enjoy Reading This Article?
Here are some more articles you might like to read next: