Skip to main content
  1. Selected publications/

Functional clusters for shape, texture, and motion encoding in macaque V2

Article info #

Authors Taekjun Kim, Rohit Kamath, Gaku Hatanaka, Tomoyuki Namima, Celeste Dylla, Wyeth Bair, Anitha Pasupathy
Publication date 2026/04/08
Journal Journal of Neuroscience
DOI https://doi.org/10.1523/JNEUROSCI.1994-25.2026

Abstract #

Macaque primary visual cortex (V1) exhibits exquisite columnar organization, while midlevel area V4 does not. Here we investigated the functional organization and representational bases of intervening area V2 in three macaques (one male, two females) with high-density Neuropixels recordings and a variety of visual stimuli—shape, texture, drifting grating, and translational motion patches. We observed dense clusters of similarly tuned neurons often spanning ~500 μm, consistent with a columnar structure. In terms of representational bases, V2 responses were largely explained by stimulus features based on local image statistics: shape tuning is well modeled by a linear combination of orientation filters, and direction selectivity is stronger with surface compared with object motion, in striking contrast to V4. Overall, our results support the progression from columns to sparse clusters as neuronal representations transform from encoding local features and feature conjunctions in V1/V2 to a high-dimensional object-based code in V4.

Figures #

Fig1. Visual stimuli and identification of the V1/V2 border #

Fig1
A, Texture stimuli. We utilized 40 naturalistic texture images categorized into eight groups (rows) based on three dimensions (C, coarseness; D, directionality; R, regularity; see texture groups legend) that influence human texture perception. Each dimension takes on two levels: C, coarse (orange) versus fine (blue); D, directional (orange) versus nondirectional (blue); and R, regular (orange) versus irregular (blue). For example, textures in the first row are coarse, directional, and regular, whereas those in the last row are fine, nondirectional, and irregular. Each texture was shown in three variations: the original, contrast-reversed, and spectrally matched noise. The fourth column of the texture groups legend represents the naturalness (N) dimension: the original and contrast-reversed variations are natural (orange) and spectrally matched noise is not (blue, data not shown).  B, Shape stimuli. One rotation of the 2D shape set from (Pasupathy and Connor, 2001) is shown. A randomly selected subset (n = 50) from rotated versions of these stimuli (n = 366, see Materials and Methods) were used to study shape responses in 17/20 sessions; in the remaining sessions (3 out of 20), a subset of 120 stimuli was used.  C, Grating stimuli. Drifting sinusoidal gratings (4 Hz, 1 s) were presented at 12 directions of motion (30° steps) and five values of spatial frequency (0.25, 0.5, 1, 2, 4 cycles/°). D, Translational motion stimuli. Elliptical noise patches with mean luminance matched to the background were displayed in sequence at seven locations centered on the RF (green circle). To ensure that the motion signal relied on non-Fourier cues, a unique noise pattern was randomly generated for each location on every trial. The long axis of the elliptical patch was oriented orthogonal to the direction of motion, which varied in increments of 45°. Five spatial displacement (dX) levels were tested while keeping motion speed constant (see Materials and Methods). The figure shown here corresponds to dX = 1/3 RF diameter.  E, The exposed cortical surface (left) and imaging of ocular dominance (right) to identify V1/V2 border from one experiment. The brighter and darker stripes are right and left eye-dominant regions in area V1, respectively. V2 recordings were made posterior to the lunate sulcus but anterior to the ocular dominance bands of V1. Note that V1-V2 segregation based on the ocular dominance pattern aligns well with a transition from denser to sparser vascular patterns from V1 to V2.

Fig2. Example recording sessions #

Fig2
A–D, Data from site 22 for texture, shape, drifting grating, and translational motion stimuli (columns) are shown. A, Each panel shows mean normalized activity (in the 0–1 range, see scale bar) of simultaneously recorded neurons (x-axis) across stimulus conditions (c-axis; see Fig. S1 for details of the normalization procedure). Texture includes 120 conditions (0–39, original; 40–79, contrast-reversed; 80–119, spectrally matched noise). Shape has 50 conditions (0–49). Drifting grating comprises 60 conditions (5 spatial frequencies × 12 directions spaced 30° apart, with 0–4 being spatial frequencies at direction 0°). Translational motion includes 40 conditions (5 spatial displacements × 8 directions spaced 45° apart, with 0–4 being displacements at direction 0°).  A horizontal stripe pattern (green and orange arrows in drifting panel) indicates that multiple nearby neurons exhibit similar visual feature tuning. B, Pairwise tuning similarity between neurons is shown, with unit IDs on both axes. Color scale indicates the strength and sign of similarity, where red and blue represent positive and negative values for the Pearson’s correlation coefficient (r). Diagonal entries were replaced with NaNs and displayed in gray. C, The x-axis represents unit IDs, and the y-axis values correspond to the distance from unit# 0, defined as the neuron recorded closest to the probe tip. Neurons with similar tuning were organized into groups using clustering analysis (see Materials and Methods). The four panels display the same neurons, but colors represent different functional clusters for each stimulus type: orange marks the largest cluster. Neurons that do not belong to any cluster are represented in gray. D, Scatterplots comparing responses of a pair of example neurons. Data are the same as depicted in the corresponding columns of the normalized response matrix in A. These two units exhibited a positive correlation for texture responses but not for other stimulus classes (see pairwise tuning similarity matrix in B).  E–H, Results from site 09. The conventions are the same as those in A–D.

Fig3. Population data: Neuronal clusters in area V2 #

Fig3
A, Results from clustering analysis (see Materials and Methods) across 24 recording sites (x-axis) for the four stimulus classes. In each recording site, individual neurons are represented by dots positioned at the depths of the corresponding contact along the probe. A depth of 0 corresponds to the location of the deepest neuron recorded in each session. Distinct clusters of neurons with similar tuning are color-coded, while neurons not assigned to any cluster appear in gray. Orange indicates the largest cluster within each site, followed by green for the second largest, and red for the smallest. Note that colors are meaningful only within a recording site; the same colors across different sites do not indicate shared feature selectivity. Sites 22 and 09 correspond to the example data shown in Figure 2C,G. B, Dependence of tuning similarity on interneuronal distance. In each panel, thin gray lines correspond to individual recording sites, with each data point representing the mean value of tuning similarity between pairs of neurons from 200 μm bins positioned every 100 μm. The black line indicates the average value across recording sites. Blue lines indicate the baseline correlation derived from shuffled neuronal pairs across different penetrations using identical stimuli. Such a baseline correction is not applicable to the shape condition, because stimulus sets were randomly selected across sessions. C, Baseline-subtracted version of the data shown in B. Red lines highlight two specific penetrations (sites 09 and 10) where RF shifts along the probe were minimal (see top panels in D). D, Shift in RF position along the length of the probe. For four representative penetrations, RF eccentricity (x-axis) is plotted against the distance along the probe (y-axis). E, We quantified the relationship between functional similarity and spatial displacements of RF by measuring the mean Z-transformed correlation across all neuronal pairs recorded within a penetration against the change in RF eccentricity (from 0.1 to 0.9 percentile) across neurons recorded along the length of the probe. We found that penetrations with larger RF shifts tended to exhibit lower functional similarity. Colors correspond to penetrations identified in D.

Fig4. Comparison between V2 and V4 #

Fig4
A, B, We assessed the local functional organization for texture stimuli by calculating, for each neuron, the average baseline-corrected Z-transformed correlation with all neuronal pairs within a 500 µm distance. The same color scale is used for both top (V2) and bottom (V4). B, Histogram of local texture-tuning similarity. Histograms show the baseline-corrected Z-transformed correlations within 500 µm for areas V2 and V4. Black bars denote neurons with statistically significant positive correlations. The median value for V2 (0.17, red line) is significantly higher than that for V4 (0.04) (Mann–Whitney U test, p < 0.001). C, D, Same comparison analyses applied to shape tuning similarity in area V2 (top) and area V4 (bottom). Baseline correction was not applicable to the shape condition, because stimulus sets were randomly selected across sessions. The median value for V2 (0.10) is significantly higher than that for V4 (0.03) (Mann–Whitney U test, p < 0.001). Note that for both shape and texture stimulus types, tuning similarity of neurons in V2 is more spatially extensive than in V4, as reflected by higher correlation values within 500 µm. In A–D, panels are based on reanalysis of previously published V4 data (Namima et al., 2025).

Fig5. Orientation energy model fit to shape responses #

Fig5
A, B, For each neuron, shape responses were predicted using an orientation energy model based on linear regression. Black dots in A and black bars in B represent neurons with statistically significant goodness-of-fit (r) as determined by an F test (p < 0.05), while gray dots and bars represent neurons that did not reach statistical significance. In most sessions, we used a shape set of 50 stimuli, but in a few sessions (3 out of 20), a larger set (n = 120) was used, which led to statistical significance even with smaller correlation values (r < 0.4). The vertical line represents the median (0.40). C, A comparison of the tuning similarity matrices derived from measured (top) versus predicted (bottom) responses from the orientation energy model. Matrices exhibit qualitatively similar structure.

Fig6. Comparing direction selectivity for drifting and translational motion #

Fig6
A, Individual neurons from seven recording sites tested with drifting motion stimuli are shown as dots and colors denote orientation-selective neurons (orange), direction selective neurons (green), and neither (gray). B, Histogram of the direction selective index (DSI) for orientation and direction selective neurons (using the same color scheme as A) computed at the optimal spatial frequency determined for each neuron. C, D, Results from 12 recording sites (the seven sites shown in A plus five others) tested with translational motion stimuli are shown using the same convention as A, B. DSI values were computed at the largest spatial displacement (dX = RF/3). E, Normalized response matrix for drifting grating stimuli (60 conditions = 12 directions × 5 spatial frequencies) from one session (site 20), following the same convention as in Figure 2A. A horizontal stripe pattern indicates that multiple nearby neurons exhibit similar visual feature tuning. F, Direction tuning curves from two example neurons from the session in E. The corresponding neurons are marked with the same colors in panel E. Error bars represent the standard error of the mean. G, Normalized response matrix for translational motion stimuli for the same session in panel E. H, Tuning curves for translational motion stimuli for the same two neurons in F. Red/blue tuning curve shapes are similar (compare panels F and H), but direction selectivity for drifting motion differs from that for translational motion. I, The optimal direction (color) and DSI value (arrow length) for each stimulus type and each neuron from three recording sites (sites 07, 09, 20) are shown. Clustered preferences for direction and orientation can be observed. J, DSI for drifting grating stimuli (x-axis) and translational motion stimuli (y-axis). The data revealed no correlation between the two measures (r = −0.02, p = 0.96). Red dotted lines indicate median DSI values for translational (0.304) and drifting (0.423) motion, respectively. K, Histogram of peak difference between drifting and translational motion tuning curves across 308 neurons that exhibited orientation selectivity for both stimulus types. The histogram reveals a bimodal distribution, indicating that the difference in preferred direction is typically either near 0 or 180°.

Fig7. Neuronal clusters with different texture selectivity #

Fig7
A, Normalized response matrix (x-axis, unit ID; y-axis, texture ID) for an example session and the temporal response profile along the four texture dimensions (right). Responses to three texture variations are shown: original (top), contrast-reversed (middle), and spectral noise (bottom). The texture dimensions, C, D, and R follow definitions in Figure 1. The naturalness dimension (N) takes on two values: original and contrast-reversed are natural (orange) while spectral noise is not (blue). In the spectral noise condition, C, D, and R dimensions are marked in green and black rather than orange and blue, since they were not included in the computations shown in B. B, For each neuron, time points with statistically significant differences in mean responses between the two levels along each texture dimension (e.g., coarse vs fine) were identified and marked as orange or blue corresponding to the greater value (see Materials and Methods). For C, D, and R dimensions, only responses to naturalistic textures were used in comparisons. Note that nearly all neurons exhibit similar texture selectivity associated with a preference for coarse, irregular, and naturalistic features. C, D, The same analyses as in A, B were performed for site 10. Nearly all neurons demonstrate similar texture selectivity, showing a preference for coarse, nondirectional, and naturalistic features. E, F, Results from site 00. Nearly all neurons demonstrate similar texture selectivity, showing a preference for coarse, directional, and naturalistic features. G, H, Results from site 20. Clusters of neurons with slightly different preferences for directionality and regularity were identified across different depths, but overall texture selectivity remained similar. I, Across all recording sites, the proportions of neurons significantly modulated by coarseness (left) were calculated from mean responses within the 0–400 ms poststimulus window, and their positions were mapped along the probe (right). Gray dots indicate neurons with no preference. J, L, The same analyses as in I for directionality (J), regularity (K), and naturalness (L) features.

Fig8. Temporal dynamics of texture selectivity and its relationship with cortical depths #

Fig8
A, All recorded V2 neurons with clear visual responses were grouped into three clusters using K-means algorithm, based on the similarity of their temporal dynamics from −100 to 200 ms relative to stimulus onset. The PSTHs in the black cluster are characterized by a fast transient peak, whereas those in the blue and green clusters exhibit stronger sustained activity, persisting 100 ms after stimulus onset. B, For each PSTH cluster, selectivity indices (i.e., area under the ROC curve; see Materials and Methods) for four texture dimensions and shape are plotted as a function of time. Selectivity for coarse, naturalistic texture features and shape emerged earlier than that for directional and regular texture features. Notably, neurons with sustained responses (blue and green clusters) exhibited stronger overall selectivity. C, For each PSTH cluster, cortical layers of individual neurons were estimated by the difference of relative LFP power between gamma (50–150 Hz) and alpha–beta (10–30 Hz) frequency ranges. Positive and negative values indicate neurons located in the superficial and deep layers, respectively, with values near 0 corresponding to layer 4. D, E, The relative proportions of neurons with statistically significant modulation for each texture attribute. Two distinct time windows were analyzed to contrast early (D, 50–100 ms poststimulus onset) and later (E, 100–150 ms poststimulus onset) neuronal activity. Neurons with significant modulation are represented by orange and blue colors, with color schemes consistent with those in Figure 7.