Handling Criteria-Driven Filtering &Sampling Across Distributed Data Partitions
Authors
Nagabhushanam Bheemisetty
Independent Researcher (US)
Article Information
DOI: 10.51583/IJLTEMAS.2026.15020000122
Subject Category: Big Data
Volume/Issue: 15/2 | Page No: 1384-1393
Publication Timeline
Submitted: 2026-03-24
Published: 2026-03-23
Abstract
This paper presents an innovative five-layered architecture for finding equitable representation across the various fragmented sections of large datasets. The architecture employs a set of customizable criteria to group together similar datasets into well-balanced unbiased sample sets; while retaining complete and accurate metadata for all file boundary elements. Important features of this architecture include: a multi-level criteria engine which is capable of processing accepted and rejected streams from a variety of sources; a stratified sampling process which effectively eliminates the presence of partition skew; and a quality assurance mechanism that provides an efficient means to capture and report performance-related metrics. Benchmark partitioning results indicate a significant shift toward overall balance ratio improvements; with improvements of approximately 1.75 to 1.12 ratio, as well as an overall reduction in volume anomaly of 87% (from ±25% down to ±3.2%), and total absence of schema drift occurred at a cost of less than 5% of the overall dataset processing costs to the author(s). By implementing this method, the processing of large datasets has been made much more reliable than previously possible, while simultaneously minimising bias and maximising the ability to provide efficient audit trial capabilities. Future enhancements to the architecture will focus on providing distributed execution capabilities in conjunction with streaming integration into enterprise data systems.
Keywords
Partitioned Datasets, Balanced Sampling, Data Unification, Criteria Engine, File-Level Fairness, Anomaly Detection
Downloads
References
1. “Balanced Sampling”, Laurent Costa, Thomas Merly-Alpa, Département des méthodes statistiques, Version no 1, diffusée le 21 juin 2017. [Google Scholar] [Crossref]
2. “A balanced sampling approach for multi-way stratification designs for small area estimation”, Piero Demetrio Falorsi and Paolo Righi, Survey Methodology, December 2008, https://www.istat.it/en/files/2016/10/Falorsi-engSURVEY_METH.pdf. [Google Scholar] [Crossref]
3. “Data Merging: Process, Challenges, and Best Practices for Combining Data from Multiple Sources”, Ehsan Elahi, November 15, 2021, https://dataladder.com/merging-data-from-multiple-sources/. [Google Scholar] [Crossref]
4. “The Speed of Now: Examples of Real-Time Processing in Action”, Wojciech Marusarz, April 17, 2023, https://nexocode.com/blog/posts/examples-of-real-time-processing/. [Google Scholar] [Crossref]
5. “Step By Step Guide: Proportional Sampling For Data Science With Python!”, Bharath K, Oct 22, 2020, https://towardsdatascience.com/step-by-step-guide-proportional-sampling-for-data-science-with-python-8b2871159ae6/. [Google Scholar] [Crossref]
6. “Rebalance Your Portfolio? You are a Market Timer and Here’s What to Consider”, Andrew Miller, March 23rd, 2017, https://alphaarchitect.com/do-you-rebalance-your-portfolio-you-are-a-market-timer/. [Google Scholar] [Crossref]
7. “R*-Grove: Balanced Spatial Partitioning for Large-Scale Datasets”, Tin Vu, Ahmed Eldawy, 28 August 2020, https://www.frontiersin.org/journals/big-data/articles/10.3389/ fdata.2020.00028/full. [Google Scholar] [Crossref]
8. “Load balancing for partition-based similarity search”, Xun Tang, Maha Alabduljalil, Xin Jin, Tao Yang, 03 July 2014, https://doi.org/10.1145/2600428.2609624. [Google Scholar] [Crossref]
9. “Incremental Partitioning for Efficient Spatial Data Analytics”, Tin Vu, Ahmed Eldawy, Vagelis Hristidis, Vassilis Tsotras, 2022, https://doi.org/10.14778/3494124.3494150. [Google Scholar] [Crossref]
10. “Effective Spatial Data Partitioning for Scalable Query Processing”, Ablimit Aji, Hoang Vo, Fusheng Wang, 3 Sep 2015, https://arxiv.org/pdf/1509.00910. [Google Scholar] [Crossref]
11. “Distributed Partitioning and Processing of Large Spatial Datasets”, Ayman I. Zeidan, 2022, https://academicworks.cuny.edu/gc_etds/4640/. [Google Scholar] [Crossref]
12. “Benchmarking data partitioning techniques in HDFS for big real spatial data”, Nikolaos Niopas, July 10, 2019, https://staff.fnwi.uva.nl/a.s.z.belloum/MSctheses/MScthesis_Nikos_Niopas.pdf. [Google Scholar] [Crossref]
13. “Efficient spatial data partitioning for distributed $$k$$ k NN joins”, Ayman Zeidan, H. Vo, 2 June 2022, https://www.semanticscholar.org/paper/Efficient-spatial-data-partitioning-for-distributed-Zeidan-Vo/549f632ea0f0d800116cc4760571d0ff7e9eaeb5. [Google Scholar] [Crossref]
14. “A Performance Study of Big Spatial Data Systems”, Md Mahbub Alam, Suprio Ray, Virendra C. Bhavsar, November 6, 2018, https://doi.org/10.1145/3282834.3282841. [Google Scholar] [Crossref]
Metrics
Views & Downloads
Similar Articles
- Drone-Based Phenotyping and its Utilization in Crop Improvement: A Review
- An IoT-Enabled Smart Healthcare Monitoring System Using Machine Learning for Early Health Risk Prediction
- Strategic Integration of Artificial Intelligence for Achieving Operational Excellence in Container Freight Stations Using Evidence from Chennai, India
- Adoption of OTT platforms: Analyzing User Behavior through the UTAUT2 Model
- Customers’ Buying Behaviour in Relation to Online Food Delivery (Evidence from National Capital Region (NCR) of India)