Experimental Results
We evaluate our HiST-VQ model on 3 publicly available skeleton-based action segmentation datasets. Namely, HuGaDB, LARa and BABEL.
Comparisons on HuGaDB and LARa
We evaluate HiST-VQ on two standard skeleton-based benchmarks, HuGaDB and LARa, and compare it with existing unsupervised methods. Across both datasets, HiST-VQ consistently achieves superior performance, demonstrating the effectiveness of its hierarchical and spatiotemporal design. Notably, the model captures fine-grained motion patterns while maintaining coherent higher-level segmentation, leading to improvements in overall segmentation quality. Furthermore, the integration of temporal modeling enhances boundary detection and reduces fragmentation in predicted segments. These results highlight the advantage of jointly learning hierarchical representations over traditional flat approaches.
Comparisons on HuGaDB and LARa. Bold and underline denote the best and second best respectively.
Comparisons on BABEL Subsets
We further evaluate HiST-VQ on multiple subsets of the BABEL dataset, which introduces significant challenges due to diverse motion patterns and long temporal sequences. HiST-VQ consistently outperforms prior methods across all subsets, demonstrating strong generalization to complex and varied data distributions. The hierarchical structure enables effective modeling of compositional actions, while temporal modeling improves the consistency of segment transitions. As a result, HiST-VQ produces more accurate and stable segmentations compared to existing approaches. These findings confirm the robustness and scalability of the proposed framework.
Comparisons on BABEL subsets. Bold and underline denote the best and second best respectively.
Segment Length Bias Comparisons
We analyze the segment length bias of HiST-VQ in comparison to existing unsupervised segmentation methods. Prior approaches often produce segments with uniform or biased lengths, limiting their ability to reflect true action durations. In contrast, HiST-VQ reduces this bias by leveraging hierarchical representations and spatiotemporal modeling. The multi-level quantization allows the model to adapt to varying action durations, while temporal cues improve alignment with motion dynamics. Consequently, HiST-VQ generates more diverse and realistic segment lengths, leading to improved overall segmentation performance.
Segment Length Bias Comparisons. Bold and underline denote the best and second best respectively.
Histograms of segment lengths on BABEL Subset-3.
Qualitative Results
Qualitative results on HuGaDB, LARa, BABAL Subset-2, and BABEL Subset-3. Colored segments represent predicted actions with a particular color denoting the same action across all models.