What ARGIRA is
ARGIRA is a research and accessibility initiative that explores new ways of representing information through technology, including sonification approaches designed to create richer experiences for people with visual impairments.
Type to filter publications and sections on this page. Results appear below as a list of links.
Ongoing research · 2026
ARGIRA
Sonification
and Perceptual
Metrics
Ranero García, Jose
Zenodo · CC BY-NC 4.0
Can the visual structure of an image predict how it perceptually sounds? ARGIRA is a research project that answers this question through systematic negative results — mapping the limits of what standard visual representations can and cannot explain about sound perception.
This page is the index and summary of the ARGIRA study series (I–XII), published on Zenodo with verifiable DOIs — not a single paper, but the map of the entire research program: experiments, data, findings, and the interactive tools derived from them.
The central question
How does mapping design influence the visual information that survives in the acoustic representation?
Initial question (ARGIRA I–III): Is Δ predictable from visual features? — The experiments showed that it is not. ARGIRA IV explains why.
A map of what
Δ does not explain
ARGIRA is not a predictive model. It is a framework for structured representational elimination: a sequence of experiments designed to determine which classes of visual representation are insufficient to capture the perceptual difference between sonification mappings.
The project evaluates three hierarchical levels of representation — low-level statistics, expanded feature spaces, and proxy semantic embeddings — using robust linear regression and Random Forest with 5-fold cross-validation. The corpus includes ~86 images from two heterogeneous visual collections: post-impressionist paintings and museum landscape photographs.
Negative results are not a problem for the program — they are its mechanism. Full series · zenodo.20534327 ↗ (opens in a new tab)
Before the question,
the program.
Experiments 11–15 did not emerge in a vacuum. The project accumulated a corpus of prior observations over several months — pipelines, benchmarks, nature corpora, clustering analyses — that established the phenomena to be explained and ruled out early hypotheses. These deposits are the ground from which ARGIRA I–V grows.
Version history (v1 – v9)
Additional corpora: Unified dataset of 74 images · zenodo.20402536 ↗ · Space A=f(H,S,I) · zenodo.20388416 ↗ · Dataset B_v9 (43 works, threshold v*) · zenodo.20357570 ↗
The correlation that
started it all
Analyzing 45 works with Argira's conversion rule, a statistically robust correlation emerged between a painting's chromatic variability and the sonic complexity of its sonification.
This correlation —r = 0.8869, R² = 0.8417, p < 0.001— was not designed. It emerged from the data. It is the reason the project exists: color appears to be related to sound.
16 works shown · Correlation computed on N = 45 works (full dataset) · Zenodo v10 (opens in new tab)
This correlation is the starting point, not the conclusion. The question it opens is: what visual features explain this link? Experiments 11–15 attempt to answer it — and find that standard visual representations are not sufficient.
The sonic museum argira.eus/argira-sonification/ lets you hear this correlation in action: the 16 works ordered from lowest to highest chromaticity, from Malevich to Kandinsky.
Three levels.
The same result.
Study summary
From perceptual divergence
to structural asymmetry
ARGIRA IV re-examines the negative results of the previous phases from a different perspective. Instead of directly modeling Δ as the target variable, it separately analyzes the two mappings that compose it: OPRS and RTR. This methodological shift reveals a previously hidden structural asymmetry.
Independent mapping modeling — ARGIRA IV
The negative results of ARGIRA I–III do not indicate an absence of image–sound relationships.
The difficulty arises from directly modeling a composite variable (Δ) that combines two transformations with distinct and partially opposing sensitivities. When both mappings are studied separately, clearly differentiated predictive structures emerge.
Predictor inversion
across domains
ARGIRA V investigates whether the same visual variables predict roughness across visual and acoustic domains. The study combines five independent visual corpora (n=319 images) and one acoustic corpus (n=72 sonifications), with 391 cases analyzed in total.
lum_contrast: secondary (ρ = 0.56–0.74)
edge_density here: secondary (ρ = 0.428)
The results support a multilayer sonification architecture where structural information (edge_density) and chromatic information (hue_entropy) must be represented independently, since they contribute non-redundant information to distinct roughness domains.
Structural convergence
in predictor space
ARGIRA VI expands the canonical corpus from 187 to 227 images (96 canonical paintings by 19 identified artists, plus 131 synthetic controls), using the new Sonification Pipeline v3.4 with a three-channel orthogonal architecture: hue_entropy controls the base frequency, edge_density + fractal_D control modulation depth (predictor derived by OLS, R² = 0.640), and luminance_contrast controls decay time.
The central finding is that visually very different paintings can converge toward nearly identical acoustic states within ARGIRA's mapping architecture: 80 pairs of canonical works satisfy |Δfractal_D| < 0.005 with |Δedge_density| > 0.15, yet still produce practically equal modulation depth values. At the artist level, recognizable structural signatures emerge in the fractal_D × edge_density space —Turner consistently occupies the high-fractal/low-edge region, Caravaggio the low-fractal/mid-edge region— without using any explicit stylistic label.
DOI: 10.5281/zenodo.20553976 · Pipeline v3.4 (companion): 10.5281/zenodo.20550714
A matrix to
measure association, not assume it
ARGIRA VII introduces the Visual-Acoustic Association Matrix (VAAM), a systematic representation of how five primary visual descriptors (hue_std, hue_entropy_bits, edge_density, fractal_D, luminance_contrast) and three interaction terms associate with seven acoustic dimensions, over a canonical corpus of 117 paintings processed with a single sonification operator (M1).
Four findings organize the study. First, the visual predictors group into two functionally independent families: a chromatic family (hue_std, hue_entropy_bits) that maps preferentially to frequency-related dimensions, and a structural family (edge_density, luminance_contrast) that maps to temporal and texture dimensions. Second, hue_std and hue_entropy_bits are non-redundant observables —partial correlation shows they independently predict distinct acoustic targets. Third, fractal_D shows minimal predictive power across all acoustic dimensions (maximum S ≈ 0.27), which contrasts with its documented relevance in empirical aesthetics and suggests that this sonification architecture transmits fractal information poorly. Fourth, the roughness dimension concentrates the matrix's highest nonlinearity indices.
Two associations reach near-perfect values (S ≈ 1.000) —hue_std→odd_bias and luminance_contrast→decay_s— but are not interpreted as empirical discoveries: they arise directly from deterministic equations implemented in the sonification operator itself, and are explicitly reported as structure imposed by the operator, not as emergent properties of the corpus. Distinguishing imposed relationships from observed associations is the central methodological contribution of this repository. The question of whether visual-acoustic associations survive across structurally distinct sonification operators remains open and is the objective of ARGIRA VIII (opens in new tab) ↗. The answer: hue_std→odd_bias does not survive — it is falsified under a full architectural reorganization (M3b) and reclassified as an operator artifact, not an invariant property of the corpus.
The question is no longer what sound does an image produce?
But rather: what aspect of the image do I want the user to perceive through sound?
A mapping is a
perceptual hypothesis
ARGIRA IV reveals that no sonification algorithm translates an image: it selects which visual dimension deserves to survive in the sound. This selection is not technical, it is epistemological. The mapping's designer decides —implicitly— what the image is for the listener.
A direct consequence follows from this principle: there is no single correct sonification of an image. There are as many valid sonifications as there are visual properties considered relevant to communicate.
| Visual property | Possible mapping | What the user perceives |
|---|---|---|
| Color (Hue) | hue → frequency | Chromatic climate |
| Saturation | saturation → amplitude | Emotional intensity |
| Brightness | luminance → pitch | Lightness / darkness |
| Edges | edge density → roughness | Geometry |
| Texture | roughness → modulation | Relief |
| Symmetry | symmetry → consonance | Order |
| Entropy | entropy → spectral noise | Complexity |
| Contrast | contrast → dynamic range | Tension |
| Depth | depth → reverberation | Space |
| Motion | optical flow → tempo | Dynamism |
ARGIRA currently explores: hue, chromatic entropy, spatial roughness.
Four types of
perceptual filter
Building on the results of ARGIRA IV, it is possible to sketch a first classification of mappings according to the type of visual property they prioritize and the perceptual experience they generate in the listener.
ARGIRA IV · zenodo.20531660 (opens in new tab) ↗
Convey the geometry of the image: edges, contours, orientation, composition. The listener can "feel" the visual structure.
Translate the image's palette into acoustic dimensions. Hue, saturation, and brightness as three independent channels of chromatic information.
Map the position within the image to acoustic space. X position to panning, Y position to pitch. The listener navigates the image like a territory.
Convey the emotional or expressive character of the image — not a measurable property, but the feeling it generates.
Δ cannot be modeled robustly using the visual representations evaluated.
Increasing representational complexity —from basic statistics to expanded spaces to semantic embeddings— does not consistently improve the prediction of Δ. ARGIRA IV suggests this limitation arises because Δ combines two mappings with different visual sensitivities, hiding predictive structures that emerge when each mapping is analyzed separately.
ARGIRA II · zenodo.20526610 (opens in new tab) ↗
ARGIRA III · zenodo.20530682 (opens in new tab) ↗
| Level | Best CV R² | Status |
|---|---|---|
| Low level Exp 11–13 |
No signal | |
| Mid level Exp 14 |
Marginal | |
| Semantic Exp 15 |
No signal |
Sonification algorithms
are not neutral translators
Each mapping selects, preserves, and discards different properties of the image. The choice of mapping determines which aspects of the work survive in the acoustic representation.
From studying a correlation to proposing a theory of mapping design.
If this result is consolidated with more mappings, ARGIRA could move from analyzing Δ as a statistical variable to proposing a theoretical framework for mapping design in multimodal accessibility — a considerably broader contribution than explaining a perceptual difference between two algorithms.
Not "the sound of the image."
An acoustic interpretation of a specific property of the image.
RTR → atmosphere · color · chromatic distribution
Several ways to tell
the same project
ARGIRA is an auditable instrument, not a landing page. There is material dedicated to explaining it at different levels of depth — from a single sentence for someone new to the topic, to the formal architecture for an ICAD reviewer. All of them describe the same central idea: moving from a black box to an auditable process. This page (the publication index) is just one of those ways.
This page (the complete index of publications and audits, with verifiable DOIs) is one more way: the most exhaustive, designed for anyone who needs complete data-by-data traceability.
The ARGIRA program
on one page
A sonification research project that systematically evaluates which classes of visual representation can — and cannot — predict the perceptual difference between acoustic mappings. Five studies, all open access on Zenodo.
How does mapping design influence the visual information that survives in the acoustic representation? Which visual features predict the perceptual difference Δ = OPRS − RTR?
~86 images (post-impressionist paintings + museum photographs). Ridge Regression and Random Forest with 5-fold CV. Three levels of visual representation evaluated hierarchically. All materials reproducible on Zenodo.
| Study | Question | Result | DOI |
|---|---|---|---|
| ARGIRA I | Low-level features (8 variables) — Exp 11 | R² CV < 0 | 20524644 ↗ (opens in new tab) |
| ARGIRA II | PCA-expanded features (209 variables) — Exp 11–15 | R² CV ≈ 0.10 | 20526610 ↗ (opens in new tab) |
| ARGIRA III | LLM semantic descriptors — Claude Haiku, n=72 | R² CV = −0.162 | 20530682 ↗ (opens in new tab) |
| ARGIRA IV | OPRS and RTR modeled separately — structural asymmetry | OPRS R²=0.582 RTR R²=0.086 |
20531660 ↗ (opens in new tab) |
| ARGIRA V | Cross-corpus predictor inversion — n=391, 5 visual corpora + 1 acoustic | edge_density ≠ hue_entropy ρ≈−0.10 |
20534328 ↗ (opens in new tab) |
Δ cannot be robustly modeled using the visual representations evaluated. ARGIRA IV reveals why: OPRS and RTR preserve structurally distinct visual properties. ARGIRA V confirms that the asymmetry generalizes across domains and corpora.
Ranero García, J. (2026). ARGIRA: Sonification and Perceptual Metrics
(Complete series I–V). Zenodo.
doi.org/10.5281/zenodo.20534327 ↗ (opens in new tab)
ICAD 2026 · Poster #4302 · Barcelona · Jul 2026
Open access
to all materials
ARGIRA Series I–XII
Correlation → structural asymmetry → invariance → falsificationFormalizes a "structural observers" framework using a transition matrix, closing the series with a methodological synthesis of the entire ARGIRA program.
Applies GLCM texture analysis to the unexplained residuals of ARGIRA X (N=227), to check whether visual structure remains uncaptured. Finds a modest fit improvement (ΔR² ≈ 0.064).
Empirical spectral analysis (SVD + Procrustes) of how visual-acoustic coupling changes depending on the sonification operator used, comparing rank-1 (M1) to rank-3 (M3) configurations.
Analyzes covariance between perceptual channels (tempo_color, fractal, space) and how that covariance reorganizes — "migrates" — when the sonification operator used is changed.
Answers the question left open by ARGIRA VII: does the hue_std→odd_bias association survive a complete change of sonification architecture (N=117, M3b reorganization)? It does not survive — it is reclassified as an operator artifact rather than an invariant property of the corpus.
Introduces the Visual-Acoustic Association Matrix (VAAM), cross-referencing 5 visual descriptors with 7 acoustic dimensions over 117 paintings + 110 synthetic controls. The chromatic and structural families predict distinct, non-redundant acoustic dimensions; fractal_D has almost no predictive power. Explicitly distinguishes associations imposed by the sonification operator from real properties of the corpus.
Extends the canonical corpus to 227 works by 19 artists. Shows that visually very different paintings can converge on nearly identical acoustic states (80 pairs with large edge difference but nearly equal modulation depth), and that per-artist structural signatures emerge without using any style label.
Across 391 cases (5 visual corpora + 1 acoustic), finds a predictor inversion across domains: edge_density dominates visually (ρ 0.49–0.84), but hue_entropy dominates acoustically (ρ = 0.595). Supports a sonification architecture with independent chromatic and structural channels.
Turning point of the series: instead of modeling Δ as a combined variable, it separates the two mappings (OPRS and RTR) and models them independently. OPRS retains predictive structure (R² CV ≈ 0.58, tied to texture and roughness); RTR barely retains it (R² CV ≈ 0.25, tied to overall color). This explains why I–III found no signal: Δ was mixing two distinct visual sensitivities.
Replaces numerical features with semantic descriptors generated by an LLM (Claude Haiku) — the "highest-level" visual representation evaluated in the series. The predictive signal disappears entirely (R² CV = −0.337 in the associated experiment 15), reinforcing that the problem was not a lack of representational complexity.
Repeats the ARGIRA I test with a much larger feature space (~209: HSV histograms, Laws filters, per-quadrant statistics, Sobel) reduced with PCA. The signal improves but remains marginal and unstable (R² CV ≈ 0.10) — it does not generalize.
First experiment in the series: tests whether ~8 low-level visual features (hue, saturation, edge density, roughness, luminance contrast) predict the perceptual divergence Δ between two sonification mappings. No signal (R² CV < 0).
Replicates on a larger corpus (N=30, vs N=9–10 in v1–v3) the hue_std→spectral roughness correlation found by chance while comparing the original mono pipeline with the stereo version (Argira v23). The correlation drops from r≈0.85 to r=0.687 as the sample grows, but remains significant (p<0.001); combined with entropy it rises to R²=0.732.
Version history (v1 – v3)
Note: when the sample is expanded from N=9–10 to N=30 in v4, the correlation drops from r≈0.85 to r=0.687 (though it remains significant, p<0.001) — a digitization artifact (Black Square) is also documented, which, when excluded, raises r to 0.790.
Checks whether the relationship between hue dispersion and roughness generalizes to independent corpora, comparing three pipelines (naive, OPRS, RTR). The naive pipeline generalizes best (r ≈ 0.54–0.62), above OPRS and RTR.
Tests the robustness of the visual-acoustic correspondence by varying the roughness threshold across a corpus of 69 works. The result hierarchy remains stable in the 1000–1750 Hz range.
Evaluates whether OPRS produces stable results when random sampling is repeated. It converges in practice around 50–100 repetitions.
Analyzes RTR with 10,000 bootstrap resamples; finds r = −0.417 and confirms that the relative order of values, not just their magnitude, influences the result.
Corpus of 14 works sonified with the minimal hue→frequency rule. Unsupervised clustering correctly groups 92.9% of cases, showing that the spectral dimension (Ds) acts as an invariant even with this simplified version of the pipeline.
21 landscape photographs taken with a mobile phone. Color saturation predicts the number of emergent harmonics with a very high correlation (r = 0.9644).
Unsupervised clustering (k-means, k=4, with PCA) over 74 sonified works. Identifies a cluster dominated by Rembrandt and Signac as a corpus outlier.
RMA Line — Structural Loss Audit
What survives a deterministic transformation, domain by domainLast version of the RMA line in the static package format (documents, reports, scripts) — not the final version of the RMA line overall, which is the Knowledge Explorer v1.8 that supersedes it. Adds the Expansion Score audit, comparing Random Forest against a low-order interaction model.
Version history (v1 – v1.4)
Current and most recent version of the entire RMA line. Replaces the static package of documents and scripts with an interactive application that runs offline in the browser (Python engine via Pyodide), organized into three layers: Evidence (audited results), Research (exploratory hypotheses), and Analysis (on-demand exploration, whose results stay local to the browser session and are never automatically promoted to Evidence). v1.8.0 is a stabilization release focused on release consistency, artifact integrity, and documentation clarity — it integrates the confirmed standalone HTML build as the canonical artifact, verifies that v1.7 functionality is preserved (including the v3.4 and v3.5.7 knowledge representations) and the Research Mode separation, and updates documentation and citation metadata; there are no changes to RMA's computational methodology, the Bootstrap CI calculation, the residual R² analysis, or the statistical classification logic. It documents as an unconfirmed edge case a possible message-persistence scenario in the Bootstrap interface, reviewed via static code inspection: it could only occur under a very specific interaction-timing condition with controls locked during the transition between Bootstrap completion and interface unlock; it has not been reproduced in normal use and no code change was applied in v1.8.0.
Version history (v1.4 – v1.7)
Applies the RMA auditor to symbolic MIDI score features. The duration_std variable is confirmed on the MAESTRO corpus (R²=0.32) but does not hold on Aria-MIDI (R²=−0.26).
Audits how much structure survives when documents are processed through optical character recognition (Tesseract 5.3.4), on the FUNSD corpus, with 20 structural variables evaluated across 4 phases.
Audits structural loss when compressing 109 paintings to JPEG (quality QF=10). Color is preserved almost intact (r≥0.98) but edge density degrades more (r=0.77).
Formalizes the general audit framework: permutation test, invariance analysis, and leave-one-out/bootstrap cross-validation, with 169 of 169 tests passed. v2 fixed three reproducibility bugs and applied it both to painting sonification (ARGIRA, N=227) and to an unrelated energy-efficiency dataset (UCI, N=768), with qualitatively different results across domains.
Version history (v1 – v2)
Documents the hue_std→roughness correlation across a corpus of 74 works (r=0.923). The entry also labels "phyllotaxis" and three hypotheses (H1–H3) without developing them further — see the DOI for the full content.
Perceptual Instrument · RAF Architecture
Navigable field, continuous coupling, and trajectory memoryAccessibility · Visual characterization
Alt-text with communicated uncertainty, more recent than the RMA lineThe full accessibility line (ARGIRA v1.4.x, fractal_D validations, synthetic RMA) is documented in the Accessibility ↓ section.
Historical material · Pipelines and variants
Development prior to the numbered series, kept for traceabilityShow 21 prior development deposits (v14–v21, mapping pipelines, B_v9 dataset) · full series up to v23 featured above
45 audio files generated by pipeline v3.5, with hue_std mapped to frequency, Dv to granular texture, and tempo to spatial distribution.
Compares four hue→frequency mapping functions (log, mel, stochastic, quadratic), with correlations between r=0.878 and r=0.951, against a scanline control with r=−0.116.
Compares a harmonic pipeline (09, r=+0.951) with a non-harmonic one (07, r=+0.687).
Compares the naive pipeline (07, r=+0.687), harmonic sine (09, r=+0.951), and control scanline (08, r=−0.116), with N=30.
Control pipeline with no hue→frequency mapping; corrects FREQ_MAX from 2000 to 10000 Hz. N=26, r=−0.116 (p=0.572) against the naive pipeline's r=+0.710 (p<0.001).
Five pipelines testing log, mel, stochastic, and quadratic as hue→frequency functions, with correlations between r=+0.878 and r=+0.935, N=30.
10 synthetic images isolating hue, saturation, and spatial irregularity; contrasts two images with the same hue_std but a different number of harmonics.
Acoustic characterization of 21 synthetic images (I00–I20) along a spatial irregularity gradient, with per-level metrics.
Technical note on a stereo-width paradox: Malevich yields W=0.196 and Kandinsky W=0.058, a result inverted relative to what complexity would predict.
Empirical calibration of a 3-orthogonal-channel architecture over N=187 artworks; edge_density→mod_depth r=+0.900 (OLS), with roughness dropped for collinearity.
Spectro-temporal analysis of 72 sonified artworks, 2,160 band-speed observations, with a regime transition at 4–8 kHz where ρ goes from +0.973 to −0.22.
Dataset of 72 artworks crossing 6 frequency bands and 5 speeds, with ρ = −0.9524 between the 4–8 kHz band and hue_std.
Test of geometric invariants over 74 artworks: full inversion of the process is not possible (mean error 0.709), though cx→pan has error <0.05 and edge_density→effort r=0.759.
74 public-domain artworks with hue_std→Ds r=0.9233, R²=0.8524, validated with a permutation test (N=1000, p<0.001).
Extended, not-peer-reviewed version, N=47: VTI→Ds r=0.9289, cross-validation with Sobel filter r=0.9570.
Standalone analyzer with 6 validated metrics: hue_std→Ds r=0.9289, irregularity→Sobel r=0.9570, N=47.
Integrates Claude Vision as a complementary object-description channel, via a proxy server (Node.js/Render) with fallback if the API is unavailable.
Adds spatial audio (centroid.x→pan, centroid.y→register) across 4 perceptual dimensions with the Web Audio API; this is the active version used by the sound museum.
Dataset of N=47 artworks with r=0.9289, validated with a permutation test (N=1000).
Applies the threshold v* = 7800 − 12500·hue_std to 43 real artworks (r=0.9370) and activates the client-side haptic module in the demo.
v1 (10.5281/zenodo.20354217 ↗) was the technical note that formalized the v* = k − m·hue_std threshold model across the ecosystem's three layers (Python, JS, Android haptic extension), with no dataset of its own. v2 is the first version to apply the threshold to real data (43 artworks) and report the correlation.
Adds a complementary semantic channel that describes images aloud using the Claude API and the Web Speech API, with no server required.
Interactive page bringing together experiments 1–15 and the findings of the ARGIRA I–V series, with a mapping taxonomy, conforming to WCAG AA.
Standalone analyzer with 6 validated metrics: hue_std↔Ds r=0.9289, N=47.
Extended version following the ICAD submission (v10), N=47, r=0.9289, Sobel r=0.9570; also documents a negative result with 2D FFT.
Independent replication with the corpus extended to 74 artworks, r=0.9233, R²=0.8524.
Control pipeline with spatial architecture and no hue→frequency mapping; documents a null result (r=−0.116).
Rewrites the analyzer with modular architecture and a high-shelf filter, as a browser tool built on the Web Audio API.
Adds speed control (temporal magnifier) in the 0.5×–1.5× range, preserving correlation r>0.92.
Tactile color sonification in night mode: pixel touch→pitch, saturation→timbre, brightness→volume.
Announces the position on a 3×3 grid first, followed by the color — for example, "upper left, blue."
From sonification
to communicated uncertainty
Starting in August 2026, the ARGIRA program extends to a distinct and complementary problem: how to communicate the uncertainty of an automatic alt-text estimate to a person who cannot verify it visually. Unlike commercial screen readers, which present their descriptions as facts, ARGIRA v1.4.x explicitly exposes the confidence of each estimate and the signals that support it — using deterministic image statistics, not neural networks.
This line includes three independent validation studies on the reliability of the technical signals used (stability of fractal_D, vulnerability of origin-discrimination features, and empirical RMA validation with 1,200 controlled runs), published with the same standard of methodological honesty as the rest of the program: what the signals can support is documented, as is what they cannot.
Three blocks of changes: (1) ES/EN language selector with cross-linking between versions; (2) result navigation via individual cards with explicit focus — no longer announced automatically on insertion; adaptations for JAWS (partially verified), TalkBack (inherited from earlier versions), and VoiceOver (adapted in code, not verified in a real environment); (3) control reordering, revised error text, removal of internal language from the header visible to the user. Introduces no changes to the classification heuristic or the uncertainty calculation. The two HTML artifacts (ES/EN) are self-contained: they work locally without installation or a network connection. The author reiterates that ARGIRA remains an experimental research artifact, not a finished product or validated accessibility technology.
Version history (v1.0 – v1.4.6)
Study complementary to ARGIRA v1.4.4 investigating whether the observed behavior of the fractal_D feature remains stable under changes in scale, interpolation method, intermediate resolutions, and controlled transformations, together with an audit of the associated residual (RMA-0) — distinct from the independent RMA calibration repository (1,200 runs, DOI 10.5281/zenodo.21935413). Experimental corpus of 52 images (18 AI-generated, 14 photographs, 20 paintings). Chains seven documented phases: fractal_D stability across the full corpus, analysis of individual D curves and successive differences, Lanczos interpolation control at different scales, bilinear vs. Lanczos comparison, intermediate-resolution analysis, residual transformation analysis, and the RMA-0 audit with control for image format/origin effects. The central question is whether the fractal_D signal is a stable property of the analyzed images or whether its measured value can be substantially affected by scale, interpolation, resolution, or controlled transformations — thus examining the signal itself and the conditions under which it changes, not an isolated numerical value as independent evidence of origin. The study does not claim that fractal_D alone determines the origin of an individual image, nor does it constitute a complete system for detecting AI-generated images; the results should be interpreted within the documented corpus, implementations, transformations, and experimental conditions. A 100-file reproducibility package (99 covered by an MD5 integrity manifest) organized into scripts, CSV/JSON numerical data, figures, reports, the 52-image corpus, the materials from the format/origin control phase, the RMA-0 audit materials, and the original source of box_counting_dimension included as a methodological reference to verify the fidelity of the analyzed implementation. The results are intended to inform ARGIRA's future development, including the planned evolution toward v1.4.5, but this deposit does not itself constitute the implementation of that version — publishing the research and its eventual incorporation into the software are separate steps. Complementary to, and not overlapping with, the independent study on the vulnerability and controllability of origin-discrimination signals (DOI 10.5281/zenodo.21945633), which focuses on luminance and saturation signals rather than fractal_D.
Study complementary to ARGIRA v1.4.4 investigating whether ARGIRA-specific signals can reliably distinguish between AI-generated images, real paintings, and real photographs. Main result: the signals studied do not constitute reliable evidence of origin under the experimental conditions analyzed. fractal_D does not discriminate significantly between AI-generated images and paintings when using the real production implementation on a balanced sample; luminance_contrast + sat_mean does show statistical separation between groups, but it is vulnerable to routine brightness, contrast, and saturation edits; the production score z_combined can be shifted through controlled photo edits without the image's structural features shifting in the same way; and stress tests show within-class score variability comparable to the separation observed between classes. Chains nine studies or experimental phases with verifiable artifacts: exploratory test of fractal_D and candidate features culminating in a Ronda 5 with 20 AI-generated images versus 20 paintings using a faithful port of the real ARGIRA v3.5 production implementation; separation with luminance_contrast + sat_mean; validation with development and validation splits; cross-validation with Gemini-generated images; experimental control for lighting, brightness, contrast, and saturation; luminance_mean audit; vulnerability control with real paintings; systematic stress test with 13 perturbation families on three base images; and an extended stress test with nine base images including AI-generated, paintings, and photographs. The natural-pairing and edit-based pairing phases (10C/10A) are not included as reproducible results because no corresponding verifiable artifacts were located. The study does not establish the origin of any individual image with certainty, does not constitute a complete AI-image detection system, does not evaluate every existing generator or generation pipeline, does not evaluate machine-learning architectures as alternative approaches, nor does it establish that the observed behavior necessarily generalizes to all possible datasets, generators, or image conditions — these limitations do not invalidate the findings, but they do delimit their scope. It is an independent scientific repository, not part of the ARGIRA v1.4.4 release itself; its results are intended to inform the evolution toward v1.4.5, but the publication of this study and its implementation in the HTML are separate decisions — the scientific evidence is first established and published in a traceable way, and only afterward is it decided which specific conclusions are incorporated into that future version's interface and logic, which should cite this repository by its DOI as part of the scientific traceability of the resulting changes.
Methodological validation/calibration study carried out as part of the development of version 1.4.5, continuing the RMA (Residual Mapping Analysis) analyses used in the preceding v1.4.4. Documents an empirical validation of RMA under synthetic control conditions with known ground truth, assessing its ability to discriminate between correctly specified or signal-free data and deliberately misspecified models. Five generative conditions were tested: an external pure-noise control, a correctly specified linear condition, and three deliberately misspecified conditions (nonlinear, composite, and hierarchical); each condition was evaluated at four sample sizes (n = 51, 100, 300, and 800), with 30 independent random seeds and two targets, resulting in 1,200 final runs. Under the tested conditions, correctly specified and signal-free cases produced negative residual R² values converging toward zero as sample size increased, while misspecified conditions produced positive residual R² values with increasing separation as the sample grew; the pure-noise control produced 0% observed false positives across 240 runs under the framework's default threshold (0.05). The study provides empirical support for RMA's discriminative behavior under the synthetic conditions tested, but it does not constitute a general validation of RMA for all possible configurations, nor does it validate downstream findings that use RMA in D or in ARGIRA. The repository includes complete raw results, aggregated results, the calibration script, the block-execution wrapper, the frozen experimental protocol, README documentation, and the main visualization; the final dataset contains 1,200 rows with no duplicate experimental keys and a framework integrity hash consistent across all records. Version v1.1.0 (current, linked here) adds rma_framework.py, the reference implementation of the RMA audit pipeline used to produce these results, mistakenly omitted from the earlier v1.0.0 — a code-completeness update: the underlying data, results, and figures do not change; MANIFEST.md5 and CITATION.cff were regenerated to reflect the new file and version.
Concept DOI for this line (always resolves to the latest version): 10.5281/zenodo.21830301 · Full tool-ecosystem annex: view annex ↗