[Paper Review] Calculating Shadows with U-Nets for Urban Environments (Short Paper)
This paper proposes novel generalizations of the Jaccard index to enhance similarity quantification across diverse mathematical structures, including sets, continuous regions, densities, and scalar fields. It introduces the coincidence index—combining Jaccard and interiority indices—for stricter similarity assessment, and extends the index to multisets, functions, and multivariate dependencies, with applications in modeling, pattern recognition, and network analysis.
Quantifying the similarity between two mathematical structures or datasets constitutes a particularly interesting and useful operation in several theoretical and applied problems. Aimed at this specific objective, the Jaccard index has been extensively used in the most diverse types of problems, also motivating some respective generalizations. The present work addresses further generalizations of this index, including its modification into a coincidence index capable of accounting also for the level of relative interiority between the two compared entities, as well as respective extensions for sets in continuous vector spaces, the generalization to multiset addition, densities and generic scalar fields, as well as a means to quantify the joint interdependence between two random variables. The also interesting possibility to take into account more than two sets has also been addressed, including the description of an index capable of quantifying the level of chaining between three structures. Several of the described and suggested eneralizations have been illustrated with respect to numeric case examples. It is also posited that these indices can play an important role while analyzing and integrating datasets in modeling approaches and pattern recognition activities, including as a measurement of clusters similarity or separation and as a resource for representing and analyzing complex networks.
Motivation & Objective
- To address the limitation of the standard Jaccard index in capturing the degree to which one set is contained within another.
- To extend the Jaccard index to continuous sets, densities, and scalar fields using area and integral-based formulations.
- To generalize the Jaccard index for multiset operations and functions with positive and negative values.
- To explore the relationship between similarity indices and joint dependence of random variables.
- To develop a three-set chaining index to quantify structural relationships among multiple entities.
Proposed method
- Introduces the interiority index (overlap index) to quantify set containment, and combines it with the Jaccard index via the square root to form the coincidence index.
- Adapts the Jaccard index to continuous sets by replacing set cardinality with region area, enabling application to geometric and spatial data.
- Extends the index to multisets and real-valued functions using integrals of minimum (intersection) and maximum (union) operations over continuous domains.
- Applies the generalized Jaccard and coincidence indices to compare probability density functions, sinusoidal functions, and real-world images.
- Proposes a chaining index for three sets by using one set as a bridge, computing Jaccard similarity between the middle set and the union of its intersections with the other two.
- Visualizes the behavior of Jaccard, interiority, and coincidence indices in scatterplots relative to the identity line.
Experimental results
Research questions
- RQ1How can the Jaccard index be improved to reflect not only overlap but also the relative interiority of one set within another?
- RQ2Can the Jaccard index be generalized to continuous sets and scalar fields, such as probability densities and image intensities?
- RQ3To what extent can the Jaccard and coincidence indices be applied to quantify the joint dependence between two random variables?
- RQ4How can similarity indices be extended to handle more than two sets, particularly in terms of structural chaining?
- RQ5What are the computational and conceptual advantages of using the coincidence index over the standard Jaccard index in data modeling and pattern recognition?
Key findings
- The coincidence index, defined as the square root of the product of the Jaccard and interiority indices, provides a more selective and strict measure of similarity than the Jaccard index alone.
- For continuous sets, replacing set cardinality with region area allows direct application of the Jaccard and coincidence indices, enabling similarity quantification in geometric and spatial contexts.
- The generalized Jaccard index for multisets and scalar fields uses integrals of min and max operations over continuous domains, enabling comparison of functions such as sinusoids and real-world images.
- The coincidence index was successfully applied to compare non-normalized, signed functions (e.g., cosine and sine), demonstrating robustness beyond non-negative data.
- The chaining index for three sets, defined as J(B, (A∩B) ∪(B∩C)) × (1 − J(A,C)), yielded a value of 6/7 for a specific example, indicating strong structural linkage between A, B, and C.
- The paper demonstrates that similarity indices like Jaccard and coincidence can be used to quantify joint variation between random variables, offering an alternative to Pearson correlation in certain contexts.
Better researchstarts right now
From reading papers to final review, dramatically reduce your research time.
No credit card · Free plan available
This review was created by AI and reviewed by human editors.